Large background noise, difficulty in feature extraction, and low parameter-optimization efficiency of diagnosis models are key challenges in rolling bearing fault diagnosis. To address these issues, this paper proposes a fault diagnosis framework that combines Singular Spectrum Decomposition (SSD) with a Multi-Strategy Enhanced Cuckoo Search (MS-CS) algorithm to optimize an Extreme Learning Machine (ELM). First, the raw vibration signal is decomposed via SSD and each intrinsic component’s energy contribution is computed; components whose cumulative energy exceeds 90% are retained and reconstructed, thereby effectively suppressing noise while preserving critical fault features. Next, Multiscale Permutation Entropy (MPE) is extracted from the reconstructed signal to form a high-discriminability feature set. To overcome the traditional Cuckoo Search algorithm’s tendency to become trapped in local optima and its slow convergence, Cauchy mutation and adaptive Levy flight strategies are introduced to enhance global exploration and local exploitation. Finally, the improved MS-CS algorithm is employed to optimize the ELM’s input weights and hidden-layer biases, yielding a high-precision diagnostic model. Experimental results on benchmark bearing data demonstrate an average fault recognition rate of 96%, representing improvements of 6.67% over the conventional CS-ELM and 18% over the unoptimized ELM. These findings confirm the proposed method’s effectiveness and robustness in practical engineering applications.
Large Language models have shown a remarkable ability to “converse” with humans in a natural language across myriad topics. Despite the proliferation of these models, a deep understanding of how they work under the hood remains elusive. The core of these Generative AI models is composed of layers of neural networks that employ the Transformer architecture. This architecture learns from large amounts of training data and creates new content in response to user input. In this study, we analyze the internals of the Transformer using Information Theory. To quantify the amount of information passing through a layer, we view it as an information transmission channel and compute the capacity of the channel. The highlight of our study is that, using Information-Theoretical tools, we develop techniques to visualize on an Information plane how the Transformer encodes the relationship between words in sentences while these words are projected into a high-dimensional vector space. We use Information Geometry to analyze the high-dimensional vectors in the Transformer layer and infer relationships between words based on the length of the geodesic connecting these vector distributions on a Riemannian manifold. Our tools reveal more information about these relationships than attention scores. In this study, we also show how Information-Theoretic analysis can help in troubleshooting learning problems in the Transformer layers.
It is imperative to establish an automated system for the identification of neonates (1–28 days old) and infants (29 days–12 months old) through the utilisation of the readily accessible 500 ppi fingerprint reader. This measure is crucial in addressing the issue of newborn swapping, facilitating the identification of missing children, monitoring immunisation records, maintaining comprehensive medical history, and other related purposes. The objective of this study is to demonstrate the potential for future identification of infants using fingerprints obtained from a 500 ppi fingerprint reader by employing a fusion technique that combines multiple instances of fingerprints, specifically the left thumb and right index fingers. The fingerprints were acquired from babies who were between the ages of one day and six months at the enrolment session. The sum-score fusion algorithm was implemented. The approach mentioned above yielded verification accuracies of 73.8%, 69.05%, and 57.14% for time intervals of 1 month, 3 months, and 6 months, respectively, between the enrolment and query fingerprints.
The Electric Circuits Concept Inventory (ECCI) is a set of multiple-choice questions that measures students’ understanding of DC Circuit analysis. The topics include (i) fundamental conceptual circuit topics (ii) Circuit laws and (iii) Circuit analysis techniques. The ECCI has been quite useful for student learning. In this paper, we explore the usefulness of generative AI (GAI) tools such as ChatGPT for teaching an Electric Circuits course. Does ChatGPT understand basic electrical circuits and know Kirchoff’s laws? We present various kinds of Electric Circuits problems and diagnose how successfully ChatGPT is able or unable to solve them. We attempt to answer the question: "Can ChatGPT be a useful tool to help students learn Electric Circuits more easily?". How good are the results of such Generative AI tools ? We present ideas on how the Electric Circuits Concept inventory can be enhanced using ChatGPT for student learning. Finally, the paper will also discuss the rules and regulations universities have deployed for student academic integrity and learning in the age of AI.
Hardware acceleration is very important in the proliferation of Deep Neural Network (DNN) technologies. The Single Input Partial Product 2-D (SIPP2D) 1 convolution-based architecture introduced in [1], [2] implements the convolution operation of convolution-based DNNs efficiently by reading the input pixels once and maximizing their reuse. In this paper, we describe the methodology to extend the DNN accelerator based on SIPP2D-based architectures to incorporate multiple strides. We describe the algorithm for any allowable stride, given any (square) input and kernel sizes. We present an analysis of the frequency of reuse of each input pixel and its complementary pixels for multiple strides and as well as its theoretical performance.
The Analytical Triangular Decoupling Internal Model Control (ATDIMC) technique for $2\times 2$ systems is generalized to $n\times n$ systems ( $n\ge 2)$ with delays and right-half-plane (RHP) transmission zeros. The formulation is done by first creating a triangular closed-loop transfer function matrix corresponding to the achievement of the triangular decoupling objective of restraining inverse-response and control-loop-interaction characteristics to a single plant output. Subsequently, the corresponding multivariable internal model controller is calculated, with transfer-function approximations made using an optimization algorithm that minimizes the Integral Time-Weighted Absolute Error (ITAE) of the difference between the step responses of the original and reduced expressions. It is shown that $n$ ATDIMC designs emerge that achieve the shifting of inverse responses and interactions to a least-desired output, with delays retained for all outputs and asymptotic tracking of setpoints achieved for all $n$ outputs of each design. To mitigate the possible effect of severe interaction on the least-desired output, a modification of this formulation is performed to spread inverse-response behavior to a second output, while minimizing the interaction of that output with the initial least-desired output. Simulation results for selected $3\times 3$ and $4\times 4$ systems show the effectiveness of these propositions.
Advancement in computer-aided tools towards accurate breast cancer early prediction models has proven to be advantageous, which in turn helps to reduce the mortality rate associated with this cancer. From the literature, random forest predictor has been observed to have high accuracy in comparison to other machine learning regressors, also genetic algorithm has been observed to be a good feature selection method in data pre-processing. In a bid to improve the accuracy of breast cancer predictive models, several studies have developed hybridized genetic algorithm models for feature selection, however, the order of hybridization may not have been taken into consideration, as this can have an impact on the hybridized model’s performance. Therefore, this paper proposes several high-performing predictive models using hybridized genetic algorithm, based on other learning models, while taking into consideration the placement order of the feature selection algorithms in the hybridized models. The Wisconsin Breast Cancer dataset was used as the test bench, while filter, wrapper and embedded feature selection algorithms were used in the proposed hybridized models. The performances of proposed hybridized models were compared with those of the individual learning models, considered in this work. These models include Fisher_Score, Mutual Information Gain, Correlation Chi-square test, Coefficient, Variance, Genetic Algorithm, Lasso and Linear Regressors with L1 regularization, Ridge Regressor with L2 regularization, Tree-based methods. From the performance evaluation results, the proposed hybridized Genetic Algorithm with Fisher_Score (GA + Fisher_Score) model showed promising results, as it had an accuracy score of 99.12%, thereby out-performing other proposed hybridized genetic algorithm models considered.
Path loss is a major factor affecting the performance of wireless networks in dense urban areas. This paper investigates the path loss models in Lagos Island, Nigeria, a dense urban area with high-rise buildings and high population density. This paper presents a detailed large-scale 3D ray-tracing investigation of the Lagos Island environment. Path loss analysis was conducted using the Close-In path loss model at 700 MHz for a TR/RX height of 20/2 m. The optimal path loss prediction model for the investigated environment was compared with existing empirical models, and the results show favorable agreement. The Close-in path loss model had a better prediction accuracy with an RMSE of 0.4331 dB, the ECC-33 path loss model achieved an accuracy with the least RMSE of 0.6743 dB. The EGLI path loss prediction model showed a pessimistic performance with the highest RMSE of 2.2496 dB, followed by Hata-Okumura with 1.9606 dB and COST231 extension-to-Hata path loss model with 1.9399 dB. Network service providers can adapt the projected 4G LTE network path loss prediction model to benchmark-related wireless propagation environments. The findings of this study are important for network service providers in Lagos Island and other dense urban areas. The study provides insights into the factors that affect path loss in these areas, and it can help network service providers optimize their transmit power and improve the performance of their wireless networks in areas where 5G is still in development and pilot trials like Nigeria and other parts of developing nations.
Deep Neural Networks (DNN) which employ multiple convolutional layers have shown remarkable accuracy in image recognition, image reconstruction, audio classification and other machine learning applications. However, a theoretical framework to explain the internal mechanism of these DNNs still remains elusive. There have been attempts to use Information Theory to crack open the DNN “black box” by showing that information is squeezed through an Information Bottleneck (IB) formed by the layers of the DNN. IB analysis in the literature is mostly based on fully connected networks and the analysis of convolutional neural networks remains extremely sparse. In analyzing the IB behavior of DNNs, the inputs and outputs of each layer are vectorized and the spatial and temporal properties of the images are ignored while computing the mutual information. In this work, we analyze DNNs which consist of convolutional layers. Each convolutional kernel along with the corresponding activation function is viewed as an Information Channel (IC). We use the spatiotemporal properties of the images to compute the Mutual Information (MI) between the channel input and output and demonstrate that the DNNs generalize and learn by reducing the Shannon capacity of these ICs while maximizing the accuracy of prediction.
The hardware acceleration of Deep Neural Networks (DNN) is a highly effective and viable solution for running them on mobile devices. The power of DNNs is now available at the edge in a compact and power-efficient form factor with the aid of hardware acceleration. In this paper, we introduce an architecture that uses a generalized method called Single Partial Product 2-Dimensional Convolution (SPP2D Convolution) which calculates a 2-D convolution in a fast and expedient manner. We demonstrate that the SPP2D architecture prevents the re-fetching of input weights for the calculation of partial products, and it can calculate the output of any input size and kernel with low latency and high throughput compared to other popular techniques. SPP2D based architecture can reduce the memory access and execution time related to input reuse by at least three times in comparison with the work done in Ardakani et al. (2018) and approximately nine times that of the standard sliding window approach. We have implemented the generalized SPP2D architecture on the Xilinx KC705 Kintex-7 evaluation board to illustrate that the new SPP2D algorithm is well-suited for the hardware acceleration of DNNs. We implemented LeNet-5 and VGGNet-16 using the SPP2D architecture. We demonstrate that the SPP2D based LeNet-5 has a high throughput of 5 GOP/s and 14.8 GOP/s/W and 42 GOP/s/W for the convolution operation using the SPP2D IP. Our LeNet-5 design achieves a similar throughput to Zhou and Jiang (2015) however using $\mathbf {3.3}\times $ fewer DSPs and an even smaller memory and lookup table (LUT) footprint. The SPP2D based VGGNet-16 network has a latency of 91.3 ms which is 79%,97%, 17% and 95% less than contemporary designs respectively, while running at a low power of 298 mW which is similar to the power level of these designs. The total processing time of our design with a parallelism factor of nine is 3.93 secs and it is 70% less than that in Ardakani et al. (2018) and 24% less than that in Panchbhaiyye and Ogunfunmi (2021). The SPP2D based LeNet-5 and VGGNet-16 accelerators provide a low-latency design with reduced memory access thus leading to a low-power design. As a result, SPP2D convolution is very well suited for hardware acceleration of DNNs.
Autonomous vehicle systems (a.k.a. self-driving cars) are becoming ubiquitous and are going to be a part of our high-technology future. At Santa Clara University, we have introduced a new class for senior undergraduates/graduate students to teach the basic principles involved in autonomous vehicle systems. This paper describes the outline and main components of the course, findings from the recent offerings of the course and some lessons learned so far. We hope this can help other institutions planning to develop similar course.
This Special Issue on “Adaptive Signal Processing and Machine Learning Using Entropy and Information Theory” was birthed from observations of the recent trend in the literature [...]
Jose G. Delgado-Frias合作论文数Washington State University;School of Electrical Engineering and Computer Science2
Roberto Togneri合作论文数University of Western Australia2