The graph reliability is the probability that a connected graph remains connected after the removal of a random number of its vertices or edges. In this article, the problem addressed is identifying the topological change in a graph that leads to the greatest increase in its graph reliability. We restrict our domain to a subproblem consisting of the case in which removal occurs only on the vertices of a graph (vertex reliability), and the only topological change allowed is a single edge insertion. In this setting, we describe 6 heuristics, all previously used in the context of edge reliability or other robustness measures. Specifically, we further analyze two spectral heuristics and their theoretical motivations with respect to vertex reliability. The performance of these 6 heuristics is evaluated by a set of computational experiments with 22000 graphs of orders 10 up to 20, generated using the Erdős-Rényi, Barabási-Albert, and Watts-Strogatz models, that compared the vertex reliability of each edge insertion produced by the heuristics. From the experiments, one spectral heuristic presented a superior performance versus the others. We propose an explanation for why this spectral heuristic performed so well and how its underlying principle – the Fiedler vector – is intrinsically linked to local and global connectedness information of the graph. Additionally, we present an initial refinement of this spectral heuristic and compare it against the others in order to show potential developments using spectral measures.
An algorithm to solve the optimal univariate stratification problem is proposed that consistently outperforms competing algorithms in a range of numerical experiments. Here the number of strata L and the sample size n are pre-specified, and the algorithm seeks the boundaries for the strata that minimize the variance of the estimator for the total of the stratification variable, under exact optimal allocation of the sample to the resulting strata. The algorithm combines ideas from the Tabu Search metaheuristic with an exact method for allocating the sample to the strata. Numerical experiments were carried out with 40 populations, varying numbers of strata (6) and sample sizes (2) - yielding a total of $ 40 \times 6 \times 2 = 480 $ 40x6x2=480 scenarios. Our results were compared to those obtained by four competing algorithms, including the well-known algorithm by Kozak and two algorithms based on alternative metaheuristics. The comparison showed that our proposed algorithm outperformed all the competitors considered, producing a high percentage of good-quality solutions.
Clustering algorithms are used to partition datasets associated with various real-world applications. However, in addition to the adopted algorithm, the obtained partition depends on the data distribution. Consequently, applying a single algorithm can lead to producing a poor-quality partition. Cluster Ensemble (CE) is an alternative to produce a good quality partition, as it combines different dataset partitions into a single consensus partition. According to the literature, in general, the consensus partition is less sensitive to noise when compared to that produced by a single algorithm. This work proposes a CE algorithm (BRKGA-CE) that combines: (i) three different strategies for producing base partitions; (ii) the BRKGA metaheuristic; (iii) the mean silhouette index, and (iv) an iterative method applied in the final phase of BRKGA-CE that allocates each object into a cluster. The core idea of BRKGA-CE is to find the representative objects of each cluster in order to maximize the mean silhouette index. To evaluate BRKGA-CE, computational experiments were carried out with 20 datasets and the main algorithms from the literature, where two well-known external validation indices (& Nscr;& Mscr;& Iscr; and & Ascr;& Rscr;) and hypothesis tests were applied. As a result, BRKGA-CE presented good-quality solutions, in the three strategies compared to the other algorithms.
This paper proposes two integer linear programming formulations to solve the optimal stratification problem (OSP). In this problem, once given the number of strata and the total sample size, the cutoff points and sample allocation should be determined jointly to minimize an expression of variance and to solve the OSP corresponding to an integer nonlinear optimization problem. One proposed formulation determines the optimal cutoff points considering classical allocation methods from the literature, and a second formulation produces the global optimum regarding the determination of the cutoff points and the sample allocation to the strata. Both formulations were implemented using Gurobi solver and applied to a set of 20 datasets from the literature, with population sizes ranging from 91 to 16,057, considering six different scenarios concerning the number of strata and sample sizes. The results obtained indicate that the formulations are a good alternative for the resolution of the OSP.
O livro Educação contemporânea: interfaces entre saberes e práticas educativas - VOLUME 1, apresenta a contribuição de pesquisadores e de pesquisadoras de diversas localidades do Brasil, com o intuito de exercitar a práxis educativa acompanhando as nuances da atualidade. Como problematização apriori destaca-se: Qual o papel da educação na contemporaneidade? De que forma podem ocorrer as interfaces entre os saberes e as práticas educativas? Diante da questão que coloca o sujeito em constante movimento, o livro problematiza com temáticas atentas à ciência, à inovação e à tecnologia, sem perder de vista a arguição solidária e humana que o ato de educar está interconectado.
In sampling theory, stratification corresponds to a technique used in surveys, which allows segmenting a population into homogeneous subpopulations (strata) to produce statistics with a higher level of precision. In particular, this article proposes a heuristic to solve the univariate stratification problem – widely studied in the literature. One of its versions sets the number of strata and the precision level and seeks to determine the limits that define such strata to minimize the sample size allocated to the strata. A heuristic-based on a stochastic optimization method and an exact optimization method was developed to achieve this goal. The performance of this heuristic was evaluated through computational experiments, considering its application in various populations used in other works in the literature, based on 20 scenarios that combine different numbers of strata and levels of precision. From the analysis of the obtained results, it is possible to verify that the heuristic had a performance superior to four algorithms in the literature in more than 94% of the cases, particularly concerning the known algorithm of Lavallée–Hidiroglou.
Com o objetivo de fornecer um panorama geral da Revista de Educação Pública (REP), o trabalho traz uma análise de sua comunidade por meio das redes de colaboração e contribuições. O periódico foi criado em 1992; o acervo analisado conta com 598 artigos publicados por 861 autores. Para a realização das meta-análises foi considerado o processo de KDD e conceitos de Teoria dos Grafos. Os resultados obtidos contemplam estatísticas gerais, análises das redes de colaboração, termos em destaque e autores considerados influentes. Assim, o trabalho fornece perspectivas e contribuições deste importante acervo.
Clustering algorithms are used to partition datasets associated with variousreal-world applications. However, in addition to the adopted algorithm, thequality of the obtained partition depends on the data distribution. Consequently, applying a single algorithm can, in many cases, lead to producing a poor-quality partition (measured by some validation index). Considering this fact, Cluster Ensemble is an alternative to produce a good quality partition,since combines different dataset partitions into a single consensus partition.According to the literature, in general, the consensus partition is less sensi-tive to noise in the data and presents higher quality when compared to thatproduced by a single algorithm.In order to produce a good quality consensus partition, this work proposesa Cluster Ensemble algorithm that combines: (i) three different strategies forproducing base partitions; (ii) the BRKGA metaheuristic, and (iii) the useof the mean silhouette index. To evaluate the new algorithm, computationalexperiments were carried out with 20 databases and several algorithms fromthe literature, where two well-known external validation indices (NMI andAR) and hypothesis tests were applied. As a result, it was observed that theproposed algorithm presented good-quality solutions, in the three strategies,for different datasets compared to the main algorithms in the literature.
The operability of a network concerns its ability to remain operational, despite possible failures in its links or equipment. One may model the network through a graph to evaluate and increase this operability. Its vertices and edges correspond to the users equipment and their connections, respectively. In this article, the problem addressed is identifying the topological change in the graph that leads to a greater increase in the operability of the associated network, considering the case in which failure occurs in the network equipment only. More specifically, we propose two spectral heuristics to improve the vertex reliability in graphs through a single edge insertion. The performance these heuristics and others that are usually found in the literature are evaluated by computational experiments with 22000 graphs of orders 10 up to 20, generated using the Models Erdos-Renyi, Barabasi-Albert, and Watts-Strogatz. From the experiments, it can be observed through analysis and application of statistical test, that one of the spectral heuristics presented a superior performance in relation to the others.
Over the last decades, many researchers have studied and proposed new methods for the solution of the multivariate optimal allocation problem, which can be performed by choosing one of the following goals: (i) minimizing the weighted combination of relative variances, considering the sample size fixed or (ii) minimizing the sample size in a way that the coefficients of variation are lower or equal to the previously fixed coefficients of variation. Taking each goal into account, the present article proposes two heuristic algorithms that were developed by studying the optimization technique called biased random key genetic algorithm. The computational experiments reported in the end of this work indicate that the proposed algorithms can be a good alternative for the solution of this problem, when compared with the two methods from literature.
O presente artigo traz a proposta de avaliação de quatro variantes do índice de silhueta quanto `a sua capacidade de detectar soluções de boa qualidade para problemas de agrupamento. Neste sentido, foram realizados cinco experimentos computacionais, contemplando 51 instâncias da literatura diversificadas (dados reais e artificiais). Como medidas de dissimilaridade foram utilizadas as distâncias euclidiana e de Manhattan, além de três algoritmos clássicos de agrupamento, a saber: PAM, DBSCAN e Bisecting k-means. De modo adicional, experimentos com a Estatística de Hopkins foram realizados com o intuito de verificar a existência de tendência de agrupamentos nas instâncias reais, em que o número de grupos k não é conhecido a priori. Os resultados obtidos indicam que a variante baseada na mediana constitui-se como boa alternativa para detectar soluções de qualidade.
This paper proposes an algorithm based on VNS metaheuristcs for k-medoids clustering, which is a NP-hard optimization problem. The VNS algorithm was applied in fifty data bases (instances) with small, medium, and large sizes, considering the number of clusters between 2 and 7. The obtained results from these experiments show the effectiveness of this approach, comparing it with nine other related clustering algorithms and an optimization formulation. Furthermore, we found that our algorithm obtained the optimal solutions for the vast majority of the cases.
Com o objetivo de analisar a comunidade de Informática da Educação e suas contribuições e colaborações no cenário nacional, o presente trabalho traz os resultados que foram obtidos a partir de três acervos de grande relevância na área: da RBIE, da RENOTE e do SBIE. A base de dados que foi tomada como ponto de partida contempla 119 edições e 4.497 artigos publicados por 14.680 autores. Para organizar, consolidar e extrair informações relevantes, a partir dessa base, foi utilizado um método de descoberta de conhecimento em bases de dados. Adicionalmente, conceitos de Teoria dos Grafos foram necessários para identificar autores influentes, conforme suas posições estruturais nos grafos obtidos através dos dados das coautorias relatadas nos artigos. A partir das meta-análises consideradas, das estatísticas gerais fornecidas e das análises das redes de colaboração, foram produzidas informações sólidas, que podem ser utilizadas em trabalhos futuros e em diversas vertentes de pesquisa.
The problem of finding an optimal sample stratification has been extensively studied in the literature. In this paper, we propose a heuristic optimization method for solving the univariate optimum stratification problem to minimize the sample size for a given precision level. The method is based on the variable neighborhood search metaheuristic, which was combined with an exact method. Numerical experiments were performed over a dataset of 24 instances, and the results of the proposed algorithm were compared with two very well-known methods from the literature. Our results outperformed 94% of the considered cases. Besides, we developed an enumeration algorithm to find the optimal global solution in some populations and scenarios, which enabled us to validate our metaheuristic method. Furthermore, we find that our algorithm obtained the optimal global solutions for the vast majority of the cases.
Este artigo se propõe a avaliar cinco variantes do índice de silhueta quanto à sua capacidade de detectar soluções de boa qualidade para problemas de agrupamento. Foram realizados cinco experimentos computacionais, contemplando 51 instâncias da literatura diversificadas (dados reais e artificiais). Como medidas de dissimilaridade foram utilizadas as distâncias euclidiana e de manhattan e para os algoritmos de agrupamento, PAM, DBSCAN e Bisecting k-means. Os resultados obtidos indicam que a variante baseada na mediana constitui-se como boa alternativa para detectar soluções de qualidade.
Wladimir Rodriguez合作论文数Postgrado en Computación, Facultad de Ingeniería, Universidad de Los Andes, Mérida, Venezuela4