An algorithm to solve the optimal univariate stratification problem is proposed that consistently outperforms competing algorithms in a range of numerical experiments. Here the number of strata L and the sample size n are pre-specified, and the algorithm seeks the boundaries for the strata that minimize the variance of the estimator for the total of the stratification variable, under exact optimal allocation of the sample to the resulting strata. The algorithm combines ideas from the Tabu Search metaheuristic with an exact method for allocating the sample to the strata. Numerical experiments were carried out with 40 populations, varying numbers of strata (6) and sample sizes (2) - yielding a total of $ 40 \times 6 \times 2 = 480 $ 40x6x2=480 scenarios. Our results were compared to those obtained by four competing algorithms, including the well-known algorithm by Kozak and two algorithms based on alternative metaheuristics. The comparison showed that our proposed algorithm outperformed all the competitors considered, producing a high percentage of good-quality solutions.
Clustering algorithms are used to partition datasets associated with various real-world applications. However, in addition to the adopted algorithm, the obtained partition depends on the data distribution. Consequently, applying a single algorithm can lead to producing a poor-quality partition. Cluster Ensemble (CE) is an alternative to produce a good quality partition, as it combines different dataset partitions into a single consensus partition. According to the literature, in general, the consensus partition is less sensitive to noise when compared to that produced by a single algorithm. This work proposes a CE algorithm (BRKGA-CE) that combines: (i) three different strategies for producing base partitions; (ii) the BRKGA metaheuristic; (iii) the mean silhouette index, and (iv) an iterative method applied in the final phase of BRKGA-CE that allocates each object into a cluster. The core idea of BRKGA-CE is to find the representative objects of each cluster in order to maximize the mean silhouette index. To evaluate BRKGA-CE, computational experiments were carried out with 20 datasets and the main algorithms from the literature, where two well-known external validation indices (& Nscr;& Mscr;& Iscr; and & Ascr;& Rscr;) and hypothesis tests were applied. As a result, BRKGA-CE presented good-quality solutions, in the three strategies compared to the other algorithms.
This paper proposes two integer linear programming formulations to solve the optimal stratification problem (OSP). In this problem, once given the number of strata and the total sample size, the cutoff points and sample allocation should be determined jointly to minimize an expression of variance and to solve the OSP corresponding to an integer nonlinear optimization problem. One proposed formulation determines the optimal cutoff points considering classical allocation methods from the literature, and a second formulation produces the global optimum regarding the determination of the cutoff points and the sample allocation to the strata. Both formulations were implemented using Gurobi solver and applied to a set of 20 datasets from the literature, with population sizes ranging from 91 to 16,057, considering six different scenarios concerning the number of strata and sample sizes. The results obtained indicate that the formulations are a good alternative for the resolution of the OSP.
The last few years have been marked by the insertion of renewable technologies in the global energy matrix, such as wind and solar energy, which are considered clean energies with low environmental impact. Wind turbines, responsible for the energy conversion process, are complex equipment that are expensive and susceptible to numerous failures. Monitoring turbine components can help detect failures before they occur, reducing equipment maintenance costs. This work compares the training time of different techniques for tuning hyperparameters in supervised machine-learning models for fault detection in wind turbines. Results show the importance of data optimization during model training.
In sampling theory, stratification corresponds to a technique used in surveys, which allows segmenting a population into homogeneous subpopulations (strata) to produce statistics with a higher level of precision. In particular, this article proposes a heuristic to solve the univariate stratification problem – widely studied in the literature. One of its versions sets the number of strata and the precision level and seeks to determine the limits that define such strata to minimize the sample size allocated to the strata. A heuristic-based on a stochastic optimization method and an exact optimization method was developed to achieve this goal. The performance of this heuristic was evaluated through computational experiments, considering its application in various populations used in other works in the literature, based on 20 scenarios that combine different numbers of strata and levels of precision. From the analysis of the obtained results, it is possible to verify that the heuristic had a performance superior to four algorithms in the literature in more than 94% of the cases, particularly concerning the known algorithm of Lavallée–Hidiroglou.
Com o objetivo de fornecer um panorama geral da Revista de Educação Pública (REP), o trabalho traz uma análise de sua comunidade por meio das redes de colaboração e contribuições. O periódico foi criado em 1992; o acervo analisado conta com 598 artigos publicados por 861 autores. Para a realização das meta-análises foi considerado o processo de KDD e conceitos de Teoria dos Grafos. Os resultados obtidos contemplam estatísticas gerais, análises das redes de colaboração, termos em destaque e autores considerados influentes. Assim, o trabalho fornece perspectivas e contribuições deste importante acervo.
Nowadays, technology has become dominant in the daily lives of most people around the world. From children to older people, technology is present, helping in the most diverse daily tasks and allowing accessibility. However, many times these people are just end-users, without any incentive to the development of computational thinking (CT). With advances in technologies, the abstraction of coding, programming languages, and the hardware resources involved will become a reality. However, while we have not progressed to this stage, it is necessary to encourage the development of CT teaching from an early age. This work will present state of the art concerning teaching initiatives and tools on programming (e.g., ScratchJr), robotics (e.g., KIBO), and other playful tools (e.g., Happy Maps) for the development of CT in the early ages, specifically filling the gap of CT at the kindergarten level. This survey presents a systematic review of the literature, emphasizing computational and robotic tools used in preschool classes to develop the CT. The systematic review evaluated more than 60 papers from 2010 to December 2020, electing 31 papers and adding three papers from the qualitative stage. The paper's amount was classified in taxonomy to show CT's principal tools and initiates applied to children early. To conclude this survey, an extensive discussion about the terms and authors related to this research area is present.
Clustering algorithms are used to partition datasets associated with variousreal-world applications. However, in addition to the adopted algorithm, thequality of the obtained partition depends on the data distribution. Consequently, applying a single algorithm can, in many cases, lead to producing a poor-quality partition (measured by some validation index). Considering this fact, Cluster Ensemble is an alternative to produce a good quality partition,since combines different dataset partitions into a single consensus partition.According to the literature, in general, the consensus partition is less sensi-tive to noise in the data and presents higher quality when compared to thatproduced by a single algorithm.In order to produce a good quality consensus partition, this work proposesa Cluster Ensemble algorithm that combines: (i) three different strategies forproducing base partitions; (ii) the BRKGA metaheuristic, and (iii) the useof the mean silhouette index. To evaluate the new algorithm, computationalexperiments were carried out with 20 databases and several algorithms fromthe literature, where two well-known external validation indices (NMI andAR) and hypothesis tests were applied. As a result, it was observed that theproposed algorithm presented good-quality solutions, in the three strategies,for different datasets compared to the main algorithms in the literature.
Over the last decades, many researchers have studied and proposed new methods for the solution of the multivariate optimal allocation problem, which can be performed by choosing one of the following goals: (i) minimizing the weighted combination of relative variances, considering the sample size fixed or (ii) minimizing the sample size in a way that the coefficients of variation are lower or equal to the previously fixed coefficients of variation. Taking each goal into account, the present article proposes two heuristic algorithms that were developed by studying the optimization technique called biased random key genetic algorithm. The computational experiments reported in the end of this work indicate that the proposed algorithms can be a good alternative for the solution of this problem, when compared with the two methods from literature.
Neste artigo é realizada a estimativa de parâmetros e solução de um problema de transferência de calor, onde os resultados numéricos são comparados com dados experimentais obtidos na literatura. O modelo matemático é resolvido numericamente utilizando o Método das Diferenças Finitas (MDF) com formulações explícita e implícita. Já a estimativa dos parâmetros é feita utilizando os métodos de otimização estocástica Luus-Jaakola (LJ) e Algoritmo de Colisão de Partículas (do inglês, Particle Collision Algorithm - PCA), bem como modificações propostas nos referidos métodos. Os resultados obtidos foram satisfatórios tendo as modificações apresentado uma redução do número de avaliações da função objetivo (NAF) necessários para encontrar os parâmetros de interesse, bem como possibilitou um bom ajuste entre as temperaturas obtidas experimentalmente e resultados numéricos, motivando novas aplicações em problemas de mesma natureza.
This paper presents a biased random-key genetic algorithm for k-medoids clustering problem. A novel heuristic operator was implemented and combined with a parallelized local search procedure. Experiments were carried out with fifty literature data sets with small, medium, and large sizes, considering several numbers of clusters, showed that the proposed algorithm outperformed eight other algorithms, for example, the classics PAM and CLARA algorithms. Furthermore, with the results of a linear integer programming formulation, we found that our algorithm obtained the global optimal solutions for most cases and, despite its stochastic nature, presented stability in terms of quality of the solutions obtained and the number of generations required to produce such solutions. In addition, considering the solutions (clusterings) produced by the algorithms, a relative validation index (average silhouette) was applied, where, again, was observed that our method performed well, producing cluster with a good structure.
O presente artigo traz a proposta de avaliação de quatro variantes do índice de silhueta quanto `a sua capacidade de detectar soluções de boa qualidade para problemas de agrupamento. Neste sentido, foram realizados cinco experimentos computacionais, contemplando 51 instâncias da literatura diversificadas (dados reais e artificiais). Como medidas de dissimilaridade foram utilizadas as distâncias euclidiana e de Manhattan, além de três algoritmos clássicos de agrupamento, a saber: PAM, DBSCAN e Bisecting k-means. De modo adicional, experimentos com a Estatística de Hopkins foram realizados com o intuito de verificar a existência de tendência de agrupamentos nas instâncias reais, em que o número de grupos k não é conhecido a priori. Os resultados obtidos indicam que a variante baseada na mediana constitui-se como boa alternativa para detectar soluções de qualidade.
This paper proposes an algorithm based on VNS metaheuristcs for k-medoids clustering, which is a NP-hard optimization problem. The VNS algorithm was applied in fifty data bases (instances) with small, medium, and large sizes, considering the number of clusters between 2 and 7. The obtained results from these experiments show the effectiveness of this approach, comparing it with nine other related clustering algorithms and an optimization formulation. Furthermore, we found that our algorithm obtained the optimal solutions for the vast majority of the cases.
Em dezembro de 2019, um novo coronavírus conhecido como SARS-CoV-2, causador da COVID-19, emergiu na China e se dispersou pelo mundo, provocando uma pandemia com ares apocalípticos. Desde então, a sociedade teve que adotar novos hábitos de prevenção, como usar máscaras faciais, higienizar as mãos com álcool 70% frequentemente e manter o isolamento social. Ao mesmo tempo, pesquisadores batalharam para desenvolver vacinas que pudessem combater o novo vírus. Diante de um cenário de medo e insegurança, a obtenção de informações corretas e seguras acerca da COVID-19 é de suma importância para que as pessoas se protejam e contribuam para a diminuição do número de casos e de mortes provocados pela doença. Nesse sentido, foi criado o aplicativo Quiz COVID-19, a fim de avaliar o nível de conhecimento dos estudantes de graduação do Instituto do Noroeste Fluminense de Educação Superior (INFES), da Universidade Federal Fluminense (UFF) acerca do novo coronavírus. Setenta e um graduandos dos cursos de Ciências Naturais, Computação, Matemática (Bacharelado), Matemática (Licenciatura) e Pedagogia participaram do Quiz COVID-19. A grande maioria dos graduandos apresentou um conhecimento adequado acerca do novo coronavírus, com uma média percentual de 75% de acertos. Em relação à média percentual de acertos por curso, temos que os graduandos que cursam Pedagogia atingiram a maior média percentual, 80%, seguidos dos que cursam Ciências Naturais, 76%; Matemática (Licenciatura), 75%; Matemática (Bacharelado), 74% e Computação, 71%. O trabalho realizado com os graduandos do INFES-UFF mostra a importância da obtenção de informações corretas e seguras, imprescindíveis para o enfrentamento à pandemia e à infodemia.
Com o objetivo de analisar a comunidade de Informática da Educação e suas contribuições e colaborações no cenário nacional, o presente trabalho traz os resultados que foram obtidos a partir de três acervos de grande relevância na área: da RBIE, da RENOTE e do SBIE. A base de dados que foi tomada como ponto de partida contempla 119 edições e 4.497 artigos publicados por 14.680 autores. Para organizar, consolidar e extrair informações relevantes, a partir dessa base, foi utilizado um método de descoberta de conhecimento em bases de dados. Adicionalmente, conceitos de Teoria dos Grafos foram necessários para identificar autores influentes, conforme suas posições estruturais nos grafos obtidos através dos dados das coautorias relatadas nos artigos. A partir das meta-análises consideradas, das estatísticas gerais fornecidas e das análises das redes de colaboração, foram produzidas informações sólidas, que podem ser utilizadas em trabalhos futuros e em diversas vertentes de pesquisa.
Este artigo se propõe a avaliar cinco variantes do índice de silhueta quanto à sua capacidade de detectar soluções de boa qualidade para problemas de agrupamento. Foram realizados cinco experimentos computacionais, contemplando 51 instâncias da literatura diversificadas (dados reais e artificiais). Como medidas de dissimilaridade foram utilizadas as distâncias euclidiana e de manhattan e para os algoritmos de agrupamento, PAM, DBSCAN e Bisecting k-means. Os resultados obtidos indicam que a variante baseada na mediana constitui-se como boa alternativa para detectar soluções de qualidade.
The Traveling Salesperson Problem with Hotel Selection (TSPHS) corresponds to a variant of the classic Traveling Salesman Problem (TSP) where the salesperson must establish a route in order to visit and attend all customers and return to the point of origin. At the end of each working day, if they had not attended all customers, the salesperson must go to a hotel (and stay there overnight). The objective of TSPHS is to first minimize the number of trips and then to minimize the total time spent. The present work proposes a new heuristic algorithm called BRKGA-LS, that combines characteristics and procedures of well-established metaheuristics and methods of vehicle routing problems. Based on computational experiments, carried out with 131 instances of the literature, the global optimum was achieved in about 90% of the cases where it is known (86 out of 95 instances), and the algorithm was able to improve the results of the literature in 11 of all 36 instances. Additionally, a hypothesis test was performed to validate the results and the good performance of the BRKGA-LS in comparison to other algorithms in the literature. Therefore, the proposed algorithm is a promising alternative for solving the problem addressed.
Com o objetivo de conhecer novas perspectivas em relação ao acervo da Revista TEIAS, em destaque pela alta qualidade de seus trabalhos e de excelência confirmada pelo sistema de avaliação Qualis da CAPES, o trabalho analisa sua comunidade por meio das redes de colaborações e contribuições. Criada no ano 2000 com foco na área de conhecimento Educação, em seu 21º ano de atuação superou a marca de 1.000 artigos publicados por mais de 1.300 autores. Para a realização das meta-análises foi considerado o processo de Descoberta de Conhecimento em Bases de Dados e conceitos de Teoria dos Grafos. Os resultados obtidos possuem estatísticas gerais, análises das redes de colaboração, termos em destaque ao passar dos anos e autores considerados influentes com base em sua frequência de publicação na revista, pela quantidade de autores com quem colaboram bem como por meio do uso da medida de centralidade por intermediação. Assim, o presente trabalho fornece informações e análises importantes, que podem ser consideradas em outras pesquisas e apoiam interpretações, perspectivas e contribuições deste importante acervo.