Background: Unintentional behavioral changes brought on by the COVID-19 outbreak may have contributed to the increase in reported suicidal attempts. The coronavirus pandemic era has contributed to modifying existing domestic violence, mental health, conflict, and anxiety. Moreover, quarantine and self-isolation may have resulted in melancholy, suicidal thoughts, drug and alcohol misuse, and loneliness. Therefore, it is crucial and significant to gather data on the global prevalence of suicide and suicidal attempts throughout the pandemic. Objective: This study's objective was to evaluate the tone of tweets regarding suicide and whether or not those tweets are connected to COVID-19. Methods: Twitter is one of the most widely used channels for sharing people's thoughts in various situations. A total of 9750 tweets have been found with respect to COVID-19-related suicidal ideation and other suicides. Gathered data were pre-processed, and feature vectors were constructed in order to establish a forecast paradigm by using artificial neural networks (ANN), long short-term memory (LSTM), and support vector machine (SVM). Results: The results demonstrated that ANN outperformed SVM and LSTM in terms of classification, achieving 91.33% accuracy while also having greater recall, precision, F-measure, and minimum error values. Conclusion: The findings of this study may help to categorize peoples' suicidal thoughts successfully. The results will help to identify future suicidal incidents with the help of the proposed approach and avoid such kinds of situations from occurring.
Background: The COVID-19 virus started in 2019 and badly affected the different sectors of many countries around the world. Based on this, financial difficulties, loss of loved ones, sudden anger, relationships, family disputes, and psychological distress increased, and individuals were stalled from carrying out their lifestyle in a normal way, and some individuals were even motivated to commit suicide. Objective: It is important to reduce the number of suicides and identify the reasons for this situation. Through this research, the focus is on identifying the main topics discussed relevant to suicides during the COVID-19 pandemic. Methods: Individuals use Twitter, a social media platform, to share their ideas freely and publically. We collected 9750 primary data through Twitter API (Application Programming Interface). After preprocessing and feature extraction by TF-IDF (Term Frequency-Inverse Document Frequency), we applied the LDA (Latent Dirichlet Allocation) and Probabilistic Latent Semantic Analysis (PLSA) topic modeling algorithms to identify topics. Results: Based on the LDA results, we extracted ten different topics under the three themes, such as the impact of COVID-19, human feelings, getting support, and having awareness. Intertopic Distance Map, Most Salient Terms, and Word Clouds Visualization are used to check the results. The coherence score and perplexing value are used to measure how interpretable the extracted topics are to humans. PLSA also extracted 25 topics with their probabilities, and Kullback–Leibler (KL) divergence was used to check the results. Conclusion: We were able to gain insight into human emotions and the main motivations behind suicide attempts using the topics we extracted. Expert feedback proved that LDA results were better than PLSA. Based on that, we found the main impact of COVID-19 on human lives, how human feelings were changed positively and negatively during that period, what supporting and awareness methods people used, and what they preferred. The required measures can then be taken by those responsible authorities and individuals to prevent, reduce, and get ready for this kind of suicidal incident in the future.
A software bug is a fault in the programming of software or an application. Bugs cause problems ranging from stability to operability and are typically the result of human error during the programming process. They could be the result of a mistake or error, as well as a fault or defect. Software bugs should be discovered during the testing stage of the software development life cycle, but some may go undetected until after deployment. When addressing a bug, it is critical to consider its priority, which is determined manually. However, it was a difficult task, and making the wrong decision could lead to major software failures. Therefore, the primary goal of this study is to propose an ensemble approach for predicting bug priority levels in bug reports. We make use of Bugzilla's dataset, which includes over 25,000 bug reports. After preprocessing the data, this study applies a variety of feature extraction techniques, including Glove, Word2Vec TF-IDF, and Doc2Vec. Then, a model that primarily employs seven architectures of Convolutional Neural Network (CNN) Algorithms, including AlexNet, LeNet, VGGNet, 1DCNN, ResNet, ZF Net, and DenseNet as the basic models. The five architectures with the highest accuracy were then used in the ensemble method, which included ResNet, DenseNet, LeNet, AlexNet, and 1DCNN, with the final results determined by the majority values. The ensemble approach performed with 79.18 % of the final accuracy result. Other architectures include AlexNet 77.1 %, ZF Net 44.50 %, VGG Net 39.30 %, 1DCNN 75.44 %, ResNet 77.34 %, DenseNet 77.32 %, and LeNet 48.58 %. It was discovered that the proposed ensemble model outperformed each algorithm. Finally, when a new bug is discovered, it can be added to the proposed model, which will then determine its priority level.
Software bugs are one of the most common types of defects in software products. A bug is an error or defect in a computer program that prevents it from functioning as intended. Bug prioritization is the process of determining which bugs should be fixed first. Prioritizing bugs is essential because it helps teams focus their efforts on fixing the most critical issues first. In recent years, the majority of researchers have focused their efforts on software bug priority-level prediction since currently it has been done manually. It is a time-consuming work and the accuracy level is also low. Therefore, the primary goal of this study is to propose an approach for predicting the bug priority level using deep learning algorithms. We used the Bugzilla dataset, which included more than 25,000 bug reports. After preprocessing data, we used several feature extraction methods, such as Global Vectors for Word Representation (GloVe), Word to Vector (Word2Vec), Term Frequency-Inverse Document Frequency (TF-IDF), and Document to Vector (Doc2Vec) methods to extract the features. There are four convolutional neural network (CNN) architectures, including AlexNet, 1DCNN, ResNet, and DenseNet used to compare the best architecture for predicting the bug priority. The accuracy is provided by each separate model: AlexNet 88.62%, 1DCNN 77.81%, ResNet 76.33%, and DenseNet 96.69%. This evaluation result discovered that DenseNet outperformed other CNN architectures in accuracy, precision, f-measure, recall, and error values. Using a proposed model, a new bug can be found and can easily assign a priority level to it while reducing the job of the software developers.
Because of the rapid advancement of Artificial Intelligence (AI) and natural language processing, powerful generative AI models have been built. ChatGPT, developed by OpenAI, is one of them. Its goal is to generate human-like text responses and participate in natural language conversations. It is critical to investigate public sentiment towards this cutting-edge technology because it is critical for improving ChatGPT interactions, addressing real-world applications in a variety of industries, advancing natural language interpretation, strengthening system robustness, and resolving ethical issues. As a result, the primary goal of this research is to examine people's perspectives about ChatGPT. a key AI language model using deep learning techniques and a machine learning technique with the help of Twitter data. The research collected and preprocessed 217,622 Twitter data before extracting features using the TF-IDF, Word2Vec, Doc2Vec, and GloVe. Then, it compares three algorithms: Long Short-Term Memory (LSTM), Artificial Neural Networks (ANN) Support Vector Machines (SVM). The results were validated by training and testing data and these algorithms' performance is measured using accuracy, precision, recall, F-score, and error values. The research seeks to determine the most successful algorithm with feature extraction combination through an evaluation. The results show that LSTM with TF-IDF feature extraction outperforms other techniques by 77.7%. The study gives light on public perceptions of ChatGPT, providing insights for responsible AI use and ethical considerations for the relevant parties. Based on that they can do necessary updates. Future research could look into employing ensemble learning methodologies for improve accuracy.
Our lifestyle heavily influences the decisions we make in life. A healthy lifestyle can greatly enhance one's ability to achieve one's goals and aspirations sustainably. Unfortunately, many individuals in the modern era tend to overlook the importance of this attribute, missing out on a tremendous investment in their overall well-being. Thus, mental, and physical health, as well as related attributes, play a pivotal role in determining the value of our lives. Therefore, everyone should have the opportunity to evaluate their lifestyle status. After knowing the status of their own lifestyle they can rearrange their life in a positive direction. As a solution to this issue, the study's main objective is to predict the current status of the lifestyle using deep learning algorithms. In the context of assessing lifestyle, the analysis involves a rich dataset comprising 22 distinct features drawn from an extensive dataset of over 15,000 data points. After pre-processing, to tackle the inherent complexity and variability of this dataset, our research employed deep learning models, specifically Long Short-Term Memory (LSTM), Artificial Neural Network (ANN), Convolutional Neural Network (CNN), and SVM (Support Vector Machine). After taking the individual algorithm results, to further enhance the accuracy, precision, recall, and F-measure while reducing error rates, these models were ingeniously combined in an ensemble approach. Leveraging the weighted average method, this ensemble model surpassed the individual models, yielding superior results that could achieve 94.51% accuracy. Moreover, this ensemble strategy contributed significantly to mitigating overfitting and bias, making it a robust and reliable tool for lifestyle assessment.
Today, in every academic institution as well as the university system assessing students' performance, identifying the uniqueness of each student and finding solutions to performance problems have become challenging issues. The main purpose of the study is to predict how student performance changes as a result of their behaviours, hobbies, extracurricular activities and different university activities. This study collected data from graduates via the online and supervised machine learning algorithms used to solve the problem. After pre-processing data, classification algorithms were applied, namely Random Forest, Multi-Layer Perceptron, Support Vector Machine, Naive Bayes and Decision Tree. The results show that the Multi-Layer Perceptron is the best algorithm considering the highest accuracy and lowest error values. An ensemble learning algorithm was then applied by combining those five algorithms. The best results were obtained using it, and according to the final results, ensemble learning increases the accuracy rather than each classifier.
Most people are tending to download or purchase apps, as a result of globally spreading technologies and smartphone usability. The Google Play Store app market is one of the most famous and rapidly increasing app markets and it captured the users’ thoughts, feelings, and, opinions about the appropriate apps they used. It is helpful for new users, app developers, and app creators to gain insights into the existing audience's opinion about relevant apps. Therefore, this study mainly aims to perform sentiment analysis on Google Play Store app users’ reviews based on the 15 latest apps. We collected 33,000 user reviews and implemented a machine-learning algorithms after initially pre-processing the data and extracting features through Term Frequency—Inverse Document Frequency (TF-IDF) vectorizer tool. The Artificial Neutral Network (ANN), Long Short-Term Memory (LSTM), and Support Vector Machine (SVM) algorithms were used for the comparison of results. By applying these algorithms, users’ reviews are mainly categorized into neutral, positive, and negative in each app separately. The overall results show that LSTM outperformed both ANN, and SVM and had a greater accuracy, recall, f-measure, and lowest error rates across all apps for providing valuable insights for sentiment classification. According to the outcomes, LSTM produces the best sentiment analysis outcomes for keeping track of users' app reviews. Affirming 80–90
Stress is a mental or emotional state brought on by demanding or unavoidable circumstances, also referred to as stressors. In order to prevent any unfavorable occurrences in life, it is crucial to understand human stress levels. Sleep disturbances are related to a number of physical, mental, and social problems. This study's main objective is to investigate how human stress might be detected using machine learning algorithms based on sleep-related behaviors. The obtained dataset includes various sleep habits and stress levels. Six machine learning techniques, including Multilayer Perception (MLP), Random Forest, Support Vector Machine (SVM), Decision Trees, Naïve Bayes and Logistic Regression were utilized in the classification level after the data had been preprocessed in order to compare and obtain the most accurate results. Based on the experiment results, it can be concluded that the Naïve Bayes algorithm, when used to classify the data, can do so with 91.27% accuracy, high precision, recall, and f-measure values, as well as the lowest mean absolute error (MAE) and root mean squared error rates (RMSE). We can estimate human stress levels using the study's findings, and we can address pertinent problems as soon as possible.
Personality is becoming an interesting factor in modern culture that affects the entire life of humans. People's lifestyles have changed as a result of Covid-19 in recent years, and it is crucial to understand how employees' personalities have changed because employees were forced to transition to working remotely during the Covid-19 outbreak. It is very beneficial for managing organizations and resolving economic losses. The main objective of this study is to use machine learning algorithms to predict the personality change of employees in the post Covid-19. Data was gathered from 500 Sri Lankan employees through Google form. After preprocessing data was then classified using machine learning methods such as Naive Bayes, Support Vector Machine (SVM), Decision Tree (J48), Random Forest, Multilayer Perception (MLP) and Ensemble Learning into positive, negative and no change categories. First, the SVM method generated the best results and then ensemble learning by combining the above five algorithms gives better accuracy than the individual algorithms. It also best in precision, recall, f-measure and a low error rates. The study's conclusions will aid in employers can recognize the various employee personality types and ensure organizational performance with the organizational culture.
The majority of educational institutions around the world have switched to online learning due to the COVID-19 pandemic. Since continuing education has become important during the pandemic as well, academics and students have recognized the value of online learning to avoid their challenges. The objective of this study is to categorize peoples' opinions and determine how the community used online learning during the pandemic. A total of 13,155 tweets were collected using the Twitter API. Of these, 4486 were positive about the online learning process, 4490 were negative, and 4179 were advertising for online learning. After pre-processing the tweets, Term Frequency-Inverse Document Frequency (TF-IDF) vectorizer is used to extract the feature vectors. The data was divided into three categories using the Long Short Term Memory (LSTM) and Support Vector Machine (SVM) algorithms. Sentiment analysis is used to determine how society feels about the online learning process by analyzing positive, negative, and advertisement sentiments. According to the results, LSTM beat SVM and achieved an accuracy of 88.58%. It also achieved higher precision, recall, f-measure values, and lowest error rates for 65% of the training dataset. Based on the findings, the significance of online learning as well as the absence of technologies, the internet, and other subpar educational practices were determined. It was determined that more workable solutions were needed in order to improve online education globally.
In most countries, the tourist sector becomes one of the significant factors for economic change. When it comes to Sri Lanka, the country's economy has grown primarily due to the tourism sector. Recently, the Covid-19 pandemic phase has had a significant negative impact on Sri Lanka's tourism industry. If we could recognize these problems, we would be able to prepare the tourism sector for future pandemics and ensure its stability. Studying peoples' opinions is crucial for taking the necessary measures, and Twitter is one of the greatest places to do this. This study suggests a model for categorizing tweets about the tourism sector during Covid-19 in Sri Lanka into four categories: positive, negative, commercial, and neutral. For this study, take into account 18980 tweets that were gathered on Twitter between the 2020 and 2022 time period. After selecting the 6257 tweets by prepossessing, the study extracts the feature vectors and result compares them using three different machine learning algorithms, including Support Vector Machine (SVM), Long Short-Term Memory (LSTM), and Artificial Neural Networks (ANN). These results in a forecast paradigm for sentiment analysis of the movement of the tourism industry. To build a model, 67% of training data sets and 33% of testing data sets were used. The findings showed that LSTM, ANN and SVM accuracy was 85%, 91%, and 76% respectively. Additionally, the study takes into account the evaluation results of precision, recall, f-measure, and error values. The ANN algorithm provides the greatest results for sentiment analysis for monitoring Sri Lanka's tourism industry movements, according to the final results.
Many people are very interested in reading books. Some of them like to read books for various reasons like as a hobby or to learn something new. At present, with modern technology, people have become accustomed to reading books online in addition to physically reading books. The most important factor in identifying which books are the most successful is the award-winning. The goal of this research is to forecast how the most successful books such as award/winning books are identified based on 10 distinct attributes including publishing date, average ratings, author, genre, language, publisher, number of pages, ratings, title, and reviews. A data set was gathered for this purpose through the online community platform Kaggle. The dataset provides information about books gathered from the website 'Goodread' for the period between 2000 and 2021. Firstly, the data set has been pre-processed. For that, duplicate data, unnecessary special characters, and missing values have been removed. To rank the preprocessed data, the Waikato Environment for Knowledge Analysis (WEKA) data mining tool was utilized, and results are improved by hyperparameter tuning. The predictive model is generated using six machine learning approaches at the classification level: Random Forest, Support Vector Machine (SVM), Decision Trees, Multilayer Perceptron (MLP), Logistic Regression, and Naive Bayes. A high accuracy of 86.32% was gained with the lowest error rates in addition to high precision, recall, and f-measure values using a Naive Bayes algorithm. Based on the results, the Naive Bayes algorithm makes a reliable prediction of how to find the most successful books using the above features.
YouTube is a popular social network for the distribution of videos globally and it is a one of the popular platforms that a large number of people access daily. On the YouTube platform, there are both high-quality videos and low-quality videos. Therefore, it is important to identify high-quality videos among the videos on YouTube to gain a full advantage of them without wasting time and make predictions of the available videos on YouTube. We identified view count, like count, comment count, caption, number of subscribers, tag count, total views, total videos, avg. polarity score, duration secs, title length, and description length as the main factors to decide the video quality. The purpose of this approach is to suggest seven different classification algorithms to make predictions of the quality/integrity of YouTube videos for users. The collected data set was pre-processed as necessary by cleaning the data groups and removing unnecessary attributes by the attribute ranking. This study was conducted using seven classification algorithms Random Forest, Logistic Regression, Support Vector Machine (SVM), Decision Tree, Multilayer Perception (MLP), Naive Bayes, and ensemble learning algorithm that combined the five individual algorithms. The model used 60% training data and 40% testing data for the classification by the Waikato Environment for Knowledge Analysis (WEKA) tool. The accuracy of each algorithm, Precision, Recall, and F-Measure are considered and Random Forest is identified as the best individual algorithm with 96.89% testing accuracy and ensemble learning with 97.21% testing accuracy. Based on the results, this model enables YouTube users to identify the quality/integrity of YouTube videos and take their decisions to view those videos.
The absence of mental health issues defines mental wellness. Rather, mental health is a condition of mental well-being that enables employees to manage work, achieve their potential, learn and work effectively, and have a positive impact on their employment environment. Poor mental health of employers negatively impacts the organization in different ways. It also negatively impacts their cognitive, behavioral, emotional, social, and relational functioning. Therefore, it's essential to recognize the mental health condition without being late and apply suitable treatments. The main objective of this study is to build a model in machine learning to predict the mental health condition and necessity of treatments of employees. The employees from technical, non-technical companies and self-employed were used for this study. The collected data samples are preprocessed and analyzed using the algorithms such as Naïve Bayes, Decision Tree (J48), Support Vector Machine (SVM), Multilayer Perceptron (MLP), Random Forest, and Ensemble Learning. The 10-fold cross-validation was used to validate the results. The J48 was the best single algorithm and its accuracy level was 92.85%. The ensemble learning that combined the above five algorithms had 93.16% accuracy level. When comparing the J48 and ensemble learning, ensemble learning is the best algorithm to forecast the level of necessity of treatments for the mental health of employees.
Stress is a state of mental or emotional weariness brought on by unavoidable or challenging situations. Additionally, the caliber of sleep has a crucial role in physical and mental health. Research on stress is interesting and beneficial to society, since it a hot issue. This study’s main objective is to investigate how humans are grouped based on their stress levels. The stress level is measured based on sleeprelated activities. To estimate the amount of stress, we applied machine learning (ML) algorithms. We made use of secondary information from the Kaggle website. The dataset includes seven sleeping habits and stress levels. The data set has been divided into two categories: low and high stress. We grouped humans based on their stress levels by considering their sleeping habits. After preprocessing the data, the findings are then acquired utilizing three unsupervised ML algorithms for clustering humans based on their stress level. We used Expectation Maximum (EM) methods, Simple KMeans, and Hierarchical Clustering as three different clustering algorithms to analyze the cluster. The Simple KMeans approach successfully categorized data with high accuracy, recall, f-measure values, and precision, according to results acquired using a data mining tool. According to the evaluation results, the Simple KMeans clustering method outperformed the outcomes of the other two algorithms with the accuracy 93.25%. The model can utilize to measure human stress and the connection between human stress and sleep were determined at the conclusion of the study.
Understanding human personality is essential for natural and social engagement. Arises a significant connection between users' personalities and their behavior. Our primary goal is to identify and classify individuals' personality traits based on their behaviors. Understanding personality types can help better understand preferences and potential differences between them. This study uses users' answers based on the questionnaire on personality to automatically identify the personality type based on behaviors. After pre-processing the data, we researched many classification techniques for automated recognition, including Naive Bayes, Support Vector Machine (SVM), Multilayer Perception (MLP), Random Forest, Logistic, and Decision Tree using a 10-fold cross-validation method. The second observational study combined all the algorithms using an Ensemble Learning algorithm, where by Vote algorithm. Accuracy, precision, recall, f-measure, and error values have been used to measure the systems' performance. According to the comparison analysis, SVM outperforms (85.8%) the other five personality trait detection algorithms. But after combing the five algorithms which contain the highest accuracy by Ensemble Learning algorithm obtained the highest performance (90.5%) than the SVM algorithm and obtained the highest recall, f-measure, and precision values, and the lowest error rates. It demonstrates that an ensemble learning approach that incorporates multiple distinct algorithms may yield greater accuracy than any one of its individual algorithms separately. Our findings are helpful for understanding how to manage and make a relationship with humans by predicting their personality earlier.
With the advancement of the field of education, sufficient information needed for education and most important things has become available on the internet. Students and scholars need a variety of formal and informal documents for their education purpose but, a large amount of data make it difficult to filter useful information from the internet. Therefore, these documents need labels for students and scholars who are engage in education to use the documents efficiently. As a result, document classification helps to assign a label to formal and informal documents. Text classification according to formal and informal styles is challenging for obtaining good accuracy, as linguistics differences are rich. This article proposed a document classification method based on formal and informal styles. This experiment used 200 text documents that were collected from the web targeting main two categories formal documents and informal documents. After preprocessing the text documents extract the feature vectors using the Term Frequency-Inverse Document Frequency (TF-IDF) vectorizer, and they are converted into numerical representation for adoption to the training model. The classification used Decision Tree (J48), Random Forest, Multilayer Perception (MLP), and Support Vector Machine (SVM) and it is tested with 5 folds cross-validation. Based on the experiment results of four classification algorithms, it indicates that the proposed approach using a Random Forest algorithm can classify the data with 94.97% accuracy with height precision, recall, f-measure values, and lowest error when comparing with the other algorithms.
With the outbreak of the Corona Virus Disease (COVID-19), nearly all educational associations throughout the world have been working tirelessly to supply online education. Students with opportunities for ongoing learning ensure their well-being. This study is being conducted to learn more about real community experiences with online learning facilities during the pandemic situation and the adaptation of online learning around the world following the pandemic circumstances. The Twitter API has been used to collect tweets for this study and a suitable result was produced after pooling the tweets. Out of the 8976 tweets, 4486 were positive, whereas 4490 were negative. After completing the pre-processing process of tweets, extract the feature vectors using the Term Frequency-Inverse Document Frequency (TF-IDF) vectorizer. Then, the dataset was loaded into supervised machine learning techniques such as Support Vector Machine (SVM) and Artificial Neural Network (ANN) to construct a forecast paradigm for predicting the probability of the society using the online learning procedure. According to the results, ANN beat SVM and achieved an accuracy of 81.97% with higher precision, recall, fmeasure values, and lowest error values. The unexpected outbreak of the pandemic caused significant disruptions to students' educational practice. They have a lack of access to technology gadgets, bad internet connectivity, and improper learning conditions. This effort also identifies the peculiarities of current technical techniques knowledge? in the development of distance learning theory. Additional financing and feasible strategies were determined to be required for the development of an efficient teaching-learning procedure for the aforementioned technique in the context of education across the globe.
According to the national statistics, the Information Technology (IT) industry provides massive rapidly growing strength of the workforce in Sri Lanka. When consider the job availability of the IT industry it mostly depends on the five skills of the applier. Programming language, full-stack development, front-end development, back-end development, and databases are playing a major role in IT skills. This paper presents a recent survey on those five major skills used by Sri Lankan employees in the IT industry and a recent index of their popularity and industry trends. The internet job database was selected as the data gathering source for this study. Then Statistical Package for the Social Science (SPSS) was used for analyzing the sample dataset. Job skills examine with job opportunities in entry-level, mid-level, and senior-level. Based on the analysis, research found that what are the expectations in each level of jobs in the current IT job market. The result of this research should be useful to job seekers, IT stakeholders, and also fresh undergraduates can be obtained a crystal-clear analysis about the new trends based on programming languages and other essential requirements of the current IT sector.