
Benefiting from the rapid development of the modern new energy automobile industry, lithium-ion batteries as the core components of new energy vehicles, the demand is rising. For both industry and consumers, accurately predicting the remaining useful life of lithium-ion batteries is an urgent need to avoid accidents and reduce range anxiety. Therefore, we propose a new method to solve this problem, the ARMI-CNN-LSTM-AT model. The innovative use of pre-ordered ARIMA models to predict battery decay data enhances the data entry quality of the second stage. In the second stage, the attention mechanism was added to the output of CNN-LSTM to redistribute the weight of unbalanced input features, The network input and output through sliding Windows. In the case of 40 rounds of data training, the final average prediction result reached 0.0136 RMSE and 0.0079 MAE. We demonstrated the importance of adding predicted data to predict battery decay, and also experimentally demonstrated that the network can maintain good prediction results on different batteries.
Active noise control (ANC) is a technology that uses sound waves to reduce or eliminate unwanted ambient noise in a given environment. We approached ANC using a deep neural network consisting of convolutional and attention layers, followed by deconvolutional layers. The proposed deep learning model processes the acoustic signal samples in frames. A piece-wise non-linear function is employed to represent the secondary path and loudspeaker saturation non-linearies. To assess the efficacy of the proposed model, we conducted tests using filtered white noise across different frequency bands and various tonals from a diverse database. The proposed convolutional model with attention layers yields a normalized mean square error (NMSE) recuction of upto 3 dB more as compared to that achieved using just convolutional model. Experiment results are presented for both confined and anechoic environments, providing a comprehensive evaluation of the proposed ANC approach.
This paper presents FedCVD, a federated learning model designed for predicting cardiovascular disease (CVD) by employing logistic regression and Support Vector Machine (SVM) algorithms. FedCVD utilizes the privacy and scalability advantages offered by federated learning to facilitate collaborative model training using decentralized patient data, ensuring confidentiality. To evaluate the effectiveness of the proposed model, experiments were conducted using the 10-year risk of coronary heart disease Kaggle dataset. To address data imbalance challenges, three techniques—Random Over Sampling, Random Under Sampling, and Synthetic Minority Oversampling Technique (SMOTE)—were explored. The study demonstrates promising performance,For the federated logistic regression with SMOTE achieving an AUC value of 0.7048. Comparative analysis with a centralized logistic regression model shows competitive results, with an AUC value of 0.7081 using Random Over Sampling. For the federated SVM model, an AUC value of 0.7340 is achieved using Random Under Sampling. In comparison, a centralized machine learning approach utilizing SVM and Random Over Sampling achieves an AUC value of 0.6962. These findings highlight the effectiveness of the proposed federated learning approach, surpassing the performance of centralized machine learning models for CVD prediction.
In this article, a face recognition via thermal imaging: a comparative study of traditional and CNN-based approaches is proposed. The methodology comprises two distinct components: traditional face recognition and CNN-based face recognition. In the traditional face recognition, we employ Random Forest (RF) and Support Vector Machine Classifier (SVM) techniques. Conversely, CNN-based face recognition leverages Convolutional Neural Networks (CNN) to identify individuals. Our research involves a comprehensive evaluation conducted across different databases, including the PUCV Drunk Thermal Face (PUCV-DTF) and the UCH Thermal Temporal Face (UCH-TTF) datasets. To emulate real-life scenarios, we introduce elements such as glasses, face mask, and noise into the original thermal images during experimentation. The recognition rates of traditional and CNN-based methods for small size and middle size databases achieve 90% and 100%, respectively, under a challenging condition, wearing glasses and mask. Experimental results demonstrate the feasibility and effectiveness of our proposed method, showcasing its robustness in tackling various challenges.
This study outlines the creation of an object detection model utilizing YOLOv5 Small, designed to identify and categorize ambulances on the road, distinguishing them based on various types and characteristics. The research process included the assembly of a specific dataset of ambulance images from the Philippines, which comprised seven classes. The project made use of image augmentation, leading to a dissimilarity score of 0.1574 between the original and augmented images, signifying the successful introduction of variations. The process of hyperparameter tuning was carried out, with a batch size of 8 proving to be the most effective, resulting in a mAP@50 score of 93% for the detection and classification of ambulances and vehicles.
Software defects are often expensive to fix, especially when they are identified late in development. Packages encapsulate logical functionality and are often developed by particular teams. Package-level defect prediction provides insights into defective designs or implementations in a system early. However, there is little work studying how to build prediction models at the package level. In this paper, we develop prediction models by using seven machine-learning algorithms and code metrics. After evaluating our approach on 20 open-source projects, we have presented that we can build effective models for predicting defective packages by using an appropriate set of metrics. However, there is no single set of metrics that can be generalized across all projects. Our study demonstrates the potential for machine-learning models to enable effective package-level defect prediction. This can guide testing and quality assurance to efficiently locate and fix defects.
Coffee leaf disease (CLD) is a major threat to coffee production worldwide, causing significant economic losses for farmers. Accurate and timely estimation of the severity of CLD is crucial for implementing effective control measures. In this paper, we propose a novel approach for severity estimation of CLD using a combination of the U-Net deep learning architecture and a pixel counting mechanism. The U-Net architecture is used to generate the segmented masks of the coffee leaf images. Finally to estimate the severity of CLD, we introduce a pixel counting mechanism that quantifies the proportion of affected pixels in the segmented regions. By summing the total number of diseased pixels and dividing it by the total number of pixels in the region, we obtain a severity score that reflects the extent of CLD infections. The algorithm proposed by us achieves a five class accuracy of 89.10%.
Financial statements are cornerstones of several analyses, such as loan applications, as well as for legal firms collecting evidence and analysis. They exert a significant influence on the decisions of these institutions. Streamlining the processing of these statements, regardless of their form—be it digital or hard copies—stands as a pivotal objective for banks and similar firms. This research explores the integration of Optical Character Recognition (OCR) and generative AI for automating the extraction of crucial financial data from bank statement images. Furthermore, we design an architecture to make a generic analysis possible on multiple types of financial documents by utilizing a classification model tailored to categorize bank statement documents. This facilitates seamless data preparation for subsequent analysis or model training. Emphasizing precision and efficiency, we investigate OCR model architectures designed specifically to enhance text extraction accuracy from low-resolution bank statement images. The study evaluates two different OCR model architectures—the accuracy of FSRCNN model being the best—achieving an accuracy above 93% in OCR. Additionally, we analyze a generative AI-based Q&A chatbot to simplify analysis for novice users.
Facial expression recognition (FER) is a burgeoning field within computer vision and artificial intelligence, with significant implications for human-computer interaction and emotion analysis. Recent advancements in deep learning, particularly convolutional neural networks (CNN), have revolutionized facial expression recognition by enabling the automatic extraction of discriminative features from facial images. These breakthroughs have led to remarkable accuracy improvements in recognizing basic emotions such as happiness, anger, sadness, surprise, fear, and disgust. In this paper, we focus on the fusion of data augmentation techniques with CNN to improve facial expression recognition. By leveraging the power of deep learning, transfer learning, and augmenting the training data, we aim to enhance the model's ability to accurately classify facial expressions across a wide spectrum of emotions. The proposed approach is particularly beneficial in scenarios where the availability of labeled training data is limited. Experimental results on the CK+ dataset show improvements in accuracy compared to state-of-the-art methods.
Encryption is a valid means to safeguard the safety of images, and for color images, encryption should be performed considering the intrinsic correlation between R, G, and B components. In this paper, we propose an image encryption algorithm based on a complementary map and an iterative convolutional code. Firstly, the plain image is input into the convolutional encoders for iteration to generate the correctional secret key. Secondly, we design a new complementary map. From the test data, the new chaotic map has passed the NIST testing, exhibits a good chaotic characteristic, and has a wider range of chaotic parameters. Thirdly, global scrambling is performed on the color image to disrupt the distribution between R, G, and B. Then, a row-layer and a column-layer are randomly selected to form a set of elements to be encrypted. Finally, performing global diffusion on the image further increases the safety of our scheme. Experimental results show that our algorithm has a preferable encryption effect and elevated safety.
Deep learning models, such as YOLOv5, well-known for object detection, and U-Net, used for segmentation, are known for their respective capabilities within computer vision tasks. In this study, the researchers introduced a novel framework that uses YOLOv5 and U-Net models, in combination with frame differencing techniques, to achieve early fire detection in an indoor setting. YOLOv5 was trained on a diverse dataset consisting of fire, smoke, and non-fire scenarios, while U-Net was exclusively trained on fire data. Motion detection was then implemented using frame differencing that allowed to identify fire movements effectively. The developed framework achieved an overall accuracy of 88%, outperforming the standalone YOLOv5 model with its 81% accuracy. This improvement of 7% in detection performance was influenced by the incorporation of fire motion analysis which effectively reduced false positive results. In summary, the study presents a robust framework that significantly improves fire detection in indoor environments with the help of motion analysis alongside the used deep learning models.
Federated learning can bring significant benefits to edge IoT systems in their scalability, efficiency, and application space by increasing the amount of computing for the nodes while decreasing the amount of network traffic required. On the other hand, the rise of efficient techniques in TinyML has been a significant boon for ultra low-power IoT ML sensors. Sadly, works on federated learning either use standard networks that are too large for TinyML devices, or are applied to relatively easier tasks to compensate for the performance degradation. In this work, we model federated learning on all four of the MLPerfTiny tasks with their respective baseline models and show that applying federated learning to TinyML models causes significant performance degradation. We also show that the performance degradation is exacerbated when the number of nodes increases. Finally, to address the performance degradation without compromising the original task or increasing the computational load for the client devices, we propose adding momentum to the server-side learning optimizer and show that this significantly mitigates the performance degradation effect, again reaching MLPerfTiny standards on all tasks.
Generative Adversarial Networks (GANs) exhibit great potential in many areas. In this paper, we explore their potential in multi-step time series forecasting. To the extent of our knowledge, this task has not been extensively researched yet, possibly due to its unique challenges when trying to model the original temporal behavior of the data. We propose a model for concrete multi-step forecasting where we mix the generative power of the unsupervised GAN loss with the deterministic prediction capabilities of supervised losses. We do this in a rather simple, sequential manner that proves to be helpful for both components of the architecture. The unsupervised component does its job by offering multiple generated predictions that follow the temporal dynamics of the time series, while the supervised component acts as a prediction selector that inspects the provided outputs and creates the most accurate one. We apply this approach in the energy sector, particularly using real industry data on oil well production, as provided to us by Raisa Energy. This learning approach leverages the generative component to provide superior results to those of the supervised counterpart. The approach also stabilizes the overall training, thereby improving the results and providing a more reliable training process.
Anomaly detection of Cloud POS data plays a significant role in the management activities of the tobacco industry. Effective anomaly detection helps retailers mitigate anomalous losses and optimize business plans. However, existing research related to anomaly detection of Cloud POS data is relatively insufficient and lacks interpretable detection approaches applicable to Cloud POS data. In this paper, we focus on the interpretable anomaly detection model and offer a new hybrid approach to detect anomaly in Cloud POS data. The proposed anomaly detection scoring method is based on Lasso algorithm and WOE (Weight of Evidence) coding. This detection method allows for the calculation of anomaly detection scores for Cloud POS data, with lower scores indicating a higher likelihood of anomalies. To validate the effectiveness of the method proposed in this paper, we conduct experiments using real Cloud POS data from the tobacco retail industry in QD city. The experimental results indicate that the proposed method not only has a comparable prediction accuracy, but also shows better interpretability compared with the benchmark methods such as SVM (Support Vector Machine) and Random Forest model.
Crossover is an important process in genetic algorithms. This process will swap genes between the chromosomes of the parents. The results from the crossover process may not be better than those of the parents, which affect the result of the genetic algorithm. We are interested in considering whether we should or should not crossover by checking the results before making a decision to crossover. If the result of the crossover from checking is not better than the parents, we do not crossover, but if the results of the crossover from checking are better than the parents, we do crossover. In the results from the test with one maximum problem in different lengths, where the size of the population is 100, the size of the generation is 100, the crossover rate is 0.7, and the mutation rate is 0.10, we found that the results from crossover consideration were better than those from a simple genetic algorithm in all lengths tested. In future work, we will apply this methodology to real-world problems.
With the rapidly developing artificial intelligence, metaverse, and 5G applications, the traffic in the Data Center Network exploded in the past decade. Optical switches are implemented to forward packets to reduce the impact of such traffic proliferation. However, the lack of optical processing capabilities challenges the controller and scheduler to switch copious traffic. To mitigate the packet loss, multi-task learning LSTM algorithm is deployed to predict the proportion of traffic flows to each server in DCNs. The prediction results achieve the MSE of 6.61E-3. Compared to the traditional LSTM neural network, the proposed scheme reduces the MSE by 20.4%, which can guide the bandwidth scheduling for high utilization in DCNs.
Offshore wind turbines (OWTs) installed far from land have historically faced significant maintenance costs and loss of power generation resources due to system failures. As the era of artificial intelligence progresses, predictive and anomaly detection algorithms continue to mature. There is a prevalent trend of generating exception labels for time series data using self-determined thresholds and applies in anomaly detection. However, such an approach may compromise the integrity and robustness of the data. This paper aims to address these challenges by implementing a real-time sensor data Supervisory control data acquisition (SCADA) detection mechanism for offshore wind turbines using the Squeeze and Excitation (SE) block Auto Encoder (AESE) algorithm point-to-point data anomaly detection technique. Then, verify unsupervised results with the real label generated by the fault log. Comparative analysis with prevalent algorithms confirms the superiority and credibility of the AESE algorithm in this application.
Single image super-resolution method using neural networks has achieved remarkable strides. However, most existing works rely on the architecture of Convolutional Neural Networks (CNNs) with shared kernel and the increasing of vertical depth, as result, super-image would loss high-frequency information. In every layer of human retina exists huge of neurons and some neurons in the same layer connect each other by horizontal neurons. Inspired by the neural network of human retina, a Stereo Neural Network (SterNet) is designed for blind image super-resolution. As the basic block of SterNet, Dynamic Filter Block (DFB) performs through unshared kernel, hence SterNet is able to easily obtain more spatial features and high-frequency semantic information. To expand the network width, one DFB with unshared kernels and one RRDB with shared kernels connect in parallel to construct Dynamic Filter and Rense Residual Blocks (DFDRB) and two DFDRBs are in parallel to form a Stereo feature extraction Block (SterB). At last, several SterBs are in series into a SterNet. Extensive experiments on several benchmarks show the effectiveness of the proposed method. This indicates that simulating the structure and operation of real neural networks is beneficial for improving vision application system.
In the context of developing nations like India, traditional business to business (B2B) commerce heavily relies on the establishment of robust relationships, trust, and credit arrangements between buyers and sellers. Consequently, ecommerce enterprises frequently. Established in 2016 with a vision to revolutionize trade in India through technology, Udaan is the countrys largest business to business ecommerce platform. Udaan operates across diverse product categories, including lifestyle, electronics, home and employ telecallers to cultivate buyer relationships, streamline order placement procedures, and promote special promotions. The accurate anticipation of buyer order placement behavior emerges as a pivotal factor for attaining sustainable growth, heightening competitiveness, and optimizing the efficiency of these telecallers. To address this challenge, we have employed an ensemble approach comprising XGBoost and a modified version of Poisson Gamma model to predict customer order patterns with precision. This paper provides an in-depth exploration of the strategic fusion of machine learning and an empirical Bayesian approach, bolstered by the judicious selection of pertinent features. This innovative approach has yielded a remarkable 3 times increase in customer order rates, show casing its potential for transformative impact in the ecommerce industry.
In the past years, since 2020, the outbreak of COVID-19 has alarmed the world with the speed and its spread around the world. This raised the demand of early, accurate and automated detection system for the COVID-19 as there is a scarcity of manpower in medical field. This attracted many researches using deep learning to build COVID-19 detection model. For the diagnosis of COVID-19, computed tomography scanning are being used as more accurate, non-invasive and efficient method in real-time. In this work, we have proposed a model using six different image classification techniques of deep learning on CT scan images and compared the accuracy to find the most suitable and reliable model for transfer learning to achieve best result on ResNet50 as 97.19% training and 98.05% testing accuracy. The model will automate the process of detection of the COVID-19, leading to the advancement in the field of smart health-care.