
In mountainous places, structures are being built for residential or commercial uses without the necessary safety precautions. Every year, landslides, torrential downpours, severe snowfall, earthquakes, volcanic eruptions, and floods cause buildings to collapse. The bulk of them are found in high-hill slope areas with loose soil types, close to river flows and other sorts of water sources. Therefore, these incidents have claimed thousands of lives. This paper deals with the process of automatic identification of critical buildings (residential/commercial) located in mountainous area which are on high-hill-slope, close to river flows, having loose soil type and high variations in land elevation contours. This study uses the primary data like, built-up/residential area and water body areas which are extracted from sample land use and land cover (LULC) using image classification techniques, and another important data like slope map and land elevation contour maps which are generated from digital elevation model (DEM). In addition, the supplementary data like, river maps, soil maps and other base maps, are also collected. All the data are integrated and taken into consideration for the identification and extraction of critical residential/build-up areas using spatial data mining technique.
Timely and accurate diagnosis of hepatitis C Virus is aimed in the proposed research using a novel dataset. For this purpose, numerous experiments are conducted using various machine learning models employing preprocessing techniques like feature engineering and data augmentation along with multiple heterogeneous classifiers. In addition, to detecting the onset of the disease, the proposed method also detects the stage of the disease to comprehend the severity for an appropriate follow-up treatment to prevent further damage to the health of the patient. Each experiment comprises various combinations of feature engineering approaches along with multiple heterogeneous classifiers. It was found that the machine learning pipeline employing the feature engineering approach of recursive feature elimination with support vector classifier as the estimator and a stacking ensemble classifier provides the best score for all performance metrics with a F1-score of 0.95, accuracy of 95.2 and mean square error of 0.06.
Solar observational studies are crucial for understanding the sun's behaviour, its impact on space weather, and its influence on Earth's climate. Central to this research is sunspot data analysis, a key indicator of solar activity and magnetic field variations. The study of solar differential rotation has been fundamental, with pioneering work revealing that faster equatorial rotation influences the sun's magnetic field and activity cycle. Sunspot areas, meticulously documented by observatories like the Royal Greenwich Observatory and KoSO, have been critical for analysing long-term solar activity trends. The integration of machine learning has significantly advanced sunspot data analysis, enhancing space weather forecasting and the understanding of solar phenomena. This paper employs change point analysis on KoSO sunspot and umbra area data to detect significant shifts over time, utilising nonparametric methods for their computational efficiency. Results show deviations from normality, positive trends, and significant autocorrelation in the data. The PELT algorithm reveals several significant shifts, dividing the period into distinct segments with varying statistical characteristics. These findings align with known solar cycles and highlight the importance of advanced statistical techniques in understanding solar activity.
Data mining framework and artificial intelligence (AI) have played a key part in all decision-making scenarios. Due to the significant expenses associated with creating training and testing datasets, we need to deal with a number of issues, object recognition, classification, and semantic segmentation in images of low spatial resolution. In this paper we first reviewed the machine learning and deep learning-based model for satellite health monitoring systems. We built the deep learning model - for satellite image classification. The dataset used is Satellite Image Classification Dataset-RSI-CB256. Two variants, ResNet-12 and ResNet-18 were tested on the dataset. The ResNet-18 showed over 0.94 accuracy for 5 number of epochs and the ResNet-12 showed 0.92 accuracy for training over 10 number of epochs. The result shows that the choice of employing the ResNet CNN architecture for Satellite Image Classification is certainly better than employing other available models such as FCNN, RCNN (with F-RCNN).
Nowadays, in a world dominated by social media, the content people share can have significant effects, particularly in the domain of cryptocurrency, where investors often turn to online advice. The instability of the cryptocurrency market is well known, and some social media individuals wield considerable influence over this market through their posts. Our study focuses on categorising these influential cryptocurrency influencers based on their English tweets, with the challenge of limited data availability. Two transformer-based models: sentence transformer fine-tuning (SetFit) and distilled BERT (DistilBERT), were used to classify cryptocurrency influencers into three subtasks: profile authors based on their degree of influence, main interests, and message intent. These models were evaluated on a Twitter-based dataset from PAN2023. The results show that SetFit achieved the best performance with a 0.82 F1-score, followed closely by DistilBERT with a 0.80 F1-score.
In mountainous places, structures are being built for residential or commercial uses without the necessary safety precautions. Every year, landslides, torrential downpours, severe snowfall, earthquakes, volcanic eruptions, and floods cause buildings to collapse. The bulk of them are found in high-hill slope areas with loose soil types, close to river flows and other sorts of water sources. Therefore, these incidents have claimed thousands of lives. This paper deals with the process of automatic identification of critical buildings (residential/commercial) located in mountainous area which are on high-hill-slope, close to river flows, having loose soil type and high variations in land elevation contours. This study use the primary data like, built-up/residential area and water body areas which are extracted from sample LULC (Land Use and Land Cover) using image classification techniques, and another important data like slope map and land elevation contour maps which are generated from DEM (Digital Elevation Model). In addition, the supplementary data like, river maps, soil maps and other base maps, are also collected. All the data are integrated and taken into consideration for the identification and extraction of critical residential/build-up areas using spatial data mining technique.
This research suggests a big data classification model that uses an improved deep convolutional neural network (IDCNN) and has five phases. In the first stage, Z-score normalisation is employed for preprocessing the input data. The second phase involves processing the preprocessed data for improved class imbalance using SMOTE-ENC. Then, the subsequent phase involves extracting the collection of features, which also includes raw data and features based on correlation, entropy, and MI. Then, in the fourth phase, to guarantee appropriate feature selection, an improved recursive feature elimination (IRFE) approach is employed for the selection of features is performed using the extracted features. Finally, ensemble classification using a collection of classifiers like Bi-LSTM, SVM, RNN and IDCNN is performed depending on the features that have been chosen. The IDCNN classifier is used in this case to categorise the final result by taking Bi-LSTM, SVM and RNN output scores as input.
For the healthcare sector, the right supplier selection and order quantity allocation decisions for the healthcare sector are crucial because the healthcare sector must deliver its products and services to its patients properly and on time. However, in this sector, supplier selection and order allocation decision are still not given enough attention. For this reason, there is a significant research and application gap in the literature. In this study, first, in order to determine the annual purchasing needs of the medical equipment that are vital for a healthcare centre in Ankara, Turkiye, always-better-control vital-essential-desirable (ABC-VED) analysis were used. Then six different scenarios for determined vital equipments were created by using goal programming model with GAMS (24.1.3) program to help the decision maker improve the purchase decision process. This proposed approach increases the efficiency of the decision process by providing the decision maker with alternative decision plans.
Sentiment analysis (SA) identifies sentiments in text, reviews, tweets, audio, images, and videos. Sentiment integrates emotion and thinking, with emotions being temporary while sentiments last longer. Emotion recognition and sentiment polarity analysis are gaining popularity in natural language processing due to their ability to mine social media data. This study applies machine learning (ML) classifiers such as random forest, logistic regression, support vector machine, and decision tree to classify text and speech as positive, negative, or neutral. Additionally, it explores available sentiment analysis tools and introduces the audio text emotion and sentiment analyser (ATESA). ATESA leverages ensemble-oriented classification techniques using deep learning, specifically bidirectional long-short-term memory recurrent neural networks (Bi-LSTM-RNN). It processes text, Twitter data, and speech converted into text. Experimental results show that ATESA achieves 92% accuracy, outperforming other algorithms.
In recent years, computer vision has made significant advances, expanding its knowledge and applications in various fields. An important example is the use of this technology to improve the recognition of different types of animals. This paper proposes an intelligent surveillance system that can individually identify each animal in a specific location and clearly indicate dangerous or unsuitable areas during monitoring, ensuring the safety of both people and the animals being monitored. In this context, deep learning algorithms, such as convolutional neural networks (CNN), are used to produce machine learning models capable of detecting and identifying objects in digital images. The study utilises the you only look once (YOLO) version 8 model and achieves 99.5% accuracy in animal recognition, demonstrating its effectiveness in monitoring. Additionally, a comparison between a model trained from random weight initialisation and another based on transfer learning reveals that the latter outperforms across various metrics, showing 99.5% accuracy, 99.3% recall, 99.5% mAP50, and 77.5% mAP50-95. These results highlight the advantage of transfer learning in optimising performance.
The exponential growth of digital information necessitates robust methods for entity resolution to ensure data quality and integration across datasets. This paper presents three novel node embedding algorithms for entity resolution in graph databases: RandomDeep, refined embedding and combined embedding. RandomDeep integrates iterative deepening depth first search with deep learning to capture structural and semantic characteristics. Refined embedding enhances initial graph convolutional (GCN) embeddings through random walk-based refinement. Combined embedding merges outputs from complementary algorithms to produce versatile representations adaptable to diverse graph structures. A two-stage graph summarisation technique supports this approach: initially as a blocking method to reduce computational complexity, and later during merging to consolidate redundant nodes. Evaluation datasets (DBLP-Scholar, Amazon-Google, Cora and Yellow-Yelp) demonstrate the methods' effectiveness, with area under cover precision and recall values ranging from 0.50 to 0.97 and F-measure values between 0.67 and 0.94. These results showcase accurate, efficient entity resolution in graph databases.
Social networks have emerged as a platform for disseminating information rapidly to friends, relatives, and the public. An effective text classification strategy can improve the effectiveness of online discussion. This has been a great motivation behind text analytics research. Several text classification approaches have been developed to enhance information extraction performance and address its challenges. However, traditional text data analytics are based on limited contextual and static resources and require effective intelligent techniques for automatically extracting features from the container. To address these issues, we proposed and developed a unique context-specific multi-class data analytics architecture based on deep learning, this approach improved the performance of data analytics and mainly focused on extracting various types of information that describe several attributes to improve the online conversation. The experimental results showed that the proposed multi-class data analytics provide promising results over classification accuracy, validation accuracy, validation loss, precision, recall, and F1-measure in support of text classification for information extraction.
This paper proposes: 1) to create a prediction model for the game addiction of adolescents using six data mining algorithms; 2) to optimise the models by adjusting the parameters; 3) to create an ensemble model. Bagging and boosting algorithms were investigated for improving the models. Data were collected from eight Northern Rajabhat Universities in Thailand. The results found that bagging with neural network had shown the highest performance with an accuracy of 99.35%, followed by the boosting with neural network (99.02%), the model with the best-optimised parameters of the neural network algorithm achieved by adjusting the learning rate. The best model was used to develop a web application for predicting the gaming addiction behaviours of adolescents, which would contribute to solve the problem.