In this paper, we perform the bias-variance decomposition of the mean-square deviation of the least-mean-squares algorithm during both the transient and steady-state phases. Although this solution has been extensively studied, to the best of our knowledge, this type of analysis has not been done before explicitly in this manner. We analyze a wide range of scenarios, including cases in which the filter length is not equal to that of the optimal solution and situations in the presence of impulsive noise. The conclusions thus obtained provide novel insights into the inner workings of the algorithm, and are supported by simulations. Moreover, we conduct experiments with real-world data considering an acoustic echo cancellation application, which show that the theoretical model thus obtained may perform reasonably well even when many of the assumptions made in the analysis do not hold.
This work provides a comprehensive overview of adaptive diffusion networks, from the first papers published on the subject to state-of-the-art solutions and current challenges. These networks consist of a collection of agents that can measure and learn from streaming data locally, and cooperate to improve the overall performance. Since their inception, adaptive diffusion networks have consolidated themselves as interesting tools for distributed estimation and learning, and have spun several types of solutions for these problems. We begin by discussing the technological advances that led to their emergence, and present the many ramifications of the area. We also discuss some of the most critical limitations of these types of networks in practical situations, such as energy consumption, and show techniques that have been proposed to cope with them. Finally, simulations with real-world data are presented in order to illustrate in practice the opportunities and challenges that they pose.
In this paper, we analyze the performance of an algorithm for adaptive diffusion networks that controls the number of nodes sampled per iteration based on the estimation error. The goal of this solution is to keep the nodes sampled while the estimation error is high in magnitude, and to cease their sampling when it is sufficiently low. Our model shows that this approach can preserve the convergence rate in comparison with the case in which every node is sampled permanently, while slightly improving the steady-state performance.
In this paper, we analyze the effects of random sampling on adaptive diffusion networks. These networks consist in a collection of nodes that can measure and process data, and that can communicate with each other to pursue a common goal of estimating an unknown system. In particular, we consider in our theoretical analysis the diffusion least-mean-squares algorithm in a scenario in which the nodes are randomly sampled. Hence, each node may or may not adapt its local estimate at a certain iteration. Our model shows that a reduction in the sampling probability leads to a noticeable deterioration in the convergence rate, and, if the nodes cooperate, to a slight decrease in the steady-state Network Mean-Square Deviation (NMSD), assuming that the environment is stationary and that all other parameters of the algorithm are kept fixed. Furthermore, we also investigate the effects of the random node sampling on the network stability.
Diffusion kernel algorithms are interesting tools for distributed nonlinear estimation. However, for the sake of feasibility, it is essential in practice to restrict their computational cost and the number of communications. In this paper, we propose a censoring algorithm for adaptive kernel diffusion networks based on random Fourier features that locally adapts the number of nodes censored according to the estimation error. It presents fast convergence during the transient phase and a significant reduction in the number of censored nodes in the steady state, thus reducing the energy consumption and the computational cost mainly by decreasing the amount of communication between nodes. Simulation results show that the proposed technique can significantly decrease the computational cost with less impact on the convergence rate when compared to existing solutions.
Resumo-Neste artigo é proposto um algoritmo de amostragem voltado a redes de difusão adaptativas multitarefas clusterizadas.O algoritmo proposto busca manter os nós amostrados quando o erro quadrático em seus clusters é elevado e deixa de amostrá-los, caso contrário
We propose a normalized least mean squares algorithm with variable step size. Unlike other solutions, it has low computational cost, only three parameters that are simple to choose, and its steady-state performance can be easily predicted. Simulations show a competitive performance in comparison with other solutions, and validate our theoretical analysis.
In this paper, multilayer perceptron neural networks are used to classify cardiac arrhythmias with a distributed approach.The training data is split between networks that communicate through a certain topology, in which each network does not have access to the training data of the others.This approach is employed to ensure patient data privacy.To obtain a clinically realistic result, data from the same patient are not considered concurrently in the training and test sets.Simulations indicate that the performance obtained with distributed training and a suitable topology is similar to the one observed with the classical training.
Multilayer perceptron neural networks are used to classify geometric figures using a distributed approach.Training data are divided among neural networks that communicate through a certain topology, in which each network does not have access to the training data of the others.Simulation results indicate that the performance achieved with distributed training and a suitable topology is similar to that observed with classical training.Thus, data privacy is guaranteed without loss of performance.
Recently, we proposed a sampling algorithm for diffusion networks that locally adapts the number of nodes sampled according to the estimation error. Thus, it reduces the computational cost associated with the learning task when the error is low in magnitude, e.g., during steady state, and maintains the sampling of the nodes otherwise, which enables fast convergence during the transient. However, its performance depends crucially on the choice of the parameter responsible for penalizing the sampling, which is a function of the variance of the measurement noise across the network. Inappropriate choices affect the tracking capability of the algorithm. In this paper, we propose a different solution, which automatically adjusts its own parameters based on the noise power estimation. Although its computational cost is slightly increased, this modification removes the need for a priori knowledge of the noise variance across the network, and increases its robustness to the presence of noisy nodes in the network. Furthermore, by implementing an adaptive reset system for the sampling mechanism, we are able to significantly improve the tracking capability of the original algorithm.
Adaptive diffusion networks have attracted attention in the scientific community as an efficient solution for distributed estimation of signals. They also have been employed in nonlinear estimation problems with distributed kernel adaptive algorithms. These solutions present a high computational cost due to the dictionary of kernel algorithms. In this paper, we propose a reduced-cost kernel algorithm for adaptive diffusion networks. It is based on the Gram-Schmidt orthogonalization process, used to span the vector space of the mapped vectors contained in the dictionary, which leads to a dictionary with a reduced cardinality. By means of simulations, we observe an advantageous computational savings in comparison to classical techniques. Keywords— Adaptive diffusion networks, kernel adaptive filtering, dictionary sparsification, nonlinear signal processing. I. INTRODUÇÃO E FORMULAÇÃO DO PROBLEMA Métodos baseados em núcleo (kernel) são capazes de resolver problemas não lineares, projetando o vetor de entrada em um espaço de dimensão mais alta, onde um método linear é usado [1]. Nesses métodos, o vetor coluna u P U Ă R é mapeado em um espaço de alta dimensão H como φpuq, usando um kernel de Mercer κ : U ˆU Ñ R, tal que [2], [3] κpu,vq “ xφpuq,φpvqyH fi φpuqφpvq, (1) @u,v P U , em que p ̈qT representa o operador de transposição de uma matriz. A Eq. (1) é conhecida como truque do kernel já que o produto interno dos vetores mapeados φpuq e φpvq pode ser calculado no espaço H sem se conhecer explicitamente a função φp ̈q [1]. O kernel Gaussiano, κpu,vq “ e ́ζ}u ́v}2 , é o mais utilizado na literatura, em que } ̈ } denota a norma euclidiana, ζ “ 1{p2σ2q sendo σ a largura do kernel. Esse André A. Bueno, Daniel G. Tiglea, Renato Candido e Magno T. M. Silva, Depto. de Engenharia de Sistemas Eletrônicos, Escola Politécnica, Universidade de São Paulo, São Paulo, SP, Brasil, e-mails: {aabueno, dtiglea, renatocan, magno}@lps.usp.br. Este trabalho foi financiado pela CAPES (88887.603151/2021-00, 88887.512247/2020-00 e 001) e FAPESP (2021/02063-6). kernel induz um espaço de dimensão infinita, é numericamente estável e apresenta a capacidade de aproximação universal, por isso, costuma ser a escolha padrão [1]. A Eq. (1) possibilita gerar versões kernel de filtros adaptativos, como é o caso do algoritmo kernel least-mean-squares (KLMS) [3]. No espaço H, o algoritmo LMS é usado para atualizar o vetor coluna de coeficientes a fim de estimar o sinal desejado dpnq [3], ou seja, ωpnq“ωpn ́1q`μ ” dpnq ́φpupnqqωpn ́1q ı φpupnqq, (2) em que ωp0q“0, μ é um passo de adaptação e upnq o vetor de entrada do filtro. Como não se conhece explicitamente a função φp ̈q e o espaço H pode ser de dimensão infinita, não é possível implementar o algoritmo da forma (2). Utilizando a Eq. (1), é possível implementar o KLMS da forma [4] ωpnq“ωpn ́1q`μrdpnq ́κN pnqωpn ́1qsκN pnq, (3)
Distributed signal processing has attracted widespread attention in the scientific community due to its several advantages over centralized approaches. Recently, graph signal processing has risen to prominence, and adaptive distributed solutions have also been proposed in the area. Both in the classical framework and in graph signal processing, sampling and censoring techniques have been topics of intense research, since the cost associated with measuring, processing and/or transmitting data throughout the entire network may be prohibitive in certain applications. In this paper, we propose a low-cost adaptive mechanism for sampling and censoring over diffusion networks that uses information from more nodes when the error in the network is high and from less nodes otherwise. It presents fast convergence during transient and a significant reduction in computational cost and energy consumption in steady state. As a censoring technique, we show that it is able to noticeably outperform other solutions. We also present a theoretical analysis to give insights about its operation, and to help the choice of suitable values for its parameters.
Resumo-Recentemente, foi proposto um algoritmo de amostragem para redes de difusão que adapta localmente o número de nós amostrados segundo o erro de estimação, gerando economia energética e computacional.Seu desempenho depende da escolha do parâmetro responsável por penalizar a amostragem, que é função da potência do ruído.Escolhas inadequadas afetam a convergência após mudanças no ambiente.Neste artigo, esse parâmetro é ajustado automaticamente baseado na estimação adaptativa da potência do ruído.Embora aumente ligeiramente o custo computacional, tal modificação melhora a capacidade de rastreamento, reduz a influência de nós
In this paper, we propose a sampling mechanism for adaptive diffusion networks that adaptively changes the amount of sampled nodes based on mean-squared error in the neighborhood of each node. It presents fast convergence during transient and a significant reduction in the number of sampled nodes in steady state. Besides reducing the computational cost, the proposed mechanism can also be used as a censoring technique, thus saving energy by reducing the amount of communication between nodes. We also present a theoretical analysis to obtain lower and upper bounds for the number of network nodes sampled in steady state.
Graph signal processing has attracted attention in the signal processing community, since it is an effective tool to deal with great quantities of interrelated data. Recently, a diffusion algorithm for adaptively learning from streaming graphs signals was proposed. However, it suffers from high computational cost since all nodes in the graph are sampled even in steady state. In this paper, we propose an adaptive sampling method for this solution that allows a reduction in computational cost in steady state, while maintaining convergence rate and presenting a slightly better steady-state performance. We also present an analysis to give insights about proper choices for its adaptation parameters.
Resumo— O processamento de sinais em grafos tem atraı́do a atenção da comunidade cientı́fica por ser uma ferramenta interessante para lidar com grandes quantidades de dados interrelacionados. Recentemente, foi proposto um algoritmo difuso para a filtragem adaptativa de sinais sobre grafos. Entretanto, esse algoritmo apresenta um custo computacional elevado, pois todos os nós do grafo são amostrados mesmo em regime permanente. Neste trabalho, é proposto um método adaptativo de amostragem para esse algoritmo que permite uma redução no custo computacional em regime permanente preservando-se o desempenho do algoritmo. Também é apresentada uma análise para facilitar a escolha de seus parâmetros.