In recent years, researchers have become interested in aspect-level sentiment analysis. In the traditional sentiment analysis of documents or sentences, a label was assigned to the entire sentence or document. Whereas a sentence or document can have aspects with different sentiments. Although deep learning models have succeeded in aspect-level sentiment analysis, these models require rich labeled datasets in different domains to extract text features and sentiment analysis. This paper uses deep transfer learning for sentiment analysis of aspect-level sentiment analysis (AHDT) of social network data. The backbone of the AHDT model is a version of RoBERTa's pre-trained deep neural network specially trained to work on social data. The features extracted from the pre-trained RoBERTa network for sentiment analysis are injected into the Bi-GRU deep neural network and then the attention layer. BI-GRU can process sequences from both sides (left to right and vice versa) and extract hidden relationships. In addition, the attention layer allows the model to pay attention to the more influential aspects of the text and provide a better interpretation. Also, this article uses the Class imbalance method to balance for training the model with almost the same polarities. The test results of the AHDT model on four SemEval datasets for the aspect-sentiment analysis task show that the model has improved the F1-score value in Resturan2014, 2015, and 2016 datasets by 0.63, 27.01, and 15.93, respectively. Also, this model has increased the accuracy value in Resturan2015 and 2016 datasets to 9.21 and 0.54, respectively. In addition, the results of experimental tests in all datasets show that the obtained values of accuracy and F1-score are close to each other, which indicates the stability of the AHDT model.
Precise short-term price prediction in the highly volatile cryptocurrency market is critical for informed trading strategies. Although Temporal Fusion Transformers (TFTs) have shown potential, their direct use often struggles in the face of the market's non-stationary nature and extreme volatility. This paper introduces an adaptive TFT modeling approach leveraging dynamic subseries lengths and pattern-based categorization to enhance short-term forecasting. We propose a novel segmentation method where subseries end at relative maxima, identified when the price increase from the preceding minimum surpasses a threshold, thus capturing significant upward movements, which act as key markers for the end of a growth phase, while potentially filtering the noise. Crucially, the fixed-length pattern ending each subseries determines the category assigned to the subsequent variable-length subseries, grouping typical market responses that follow similar preceding conditions. A distinct TFT model trained for each category is specialized in predicting the evolution of these subsequent subseries based on their initial steps after the preceding peak. Experimental results on ETH-USDT 10-minute data over a two-month test period demonstrate that our adaptive approach significantly outperforms baseline fixed-length TFT and LSTM models in prediction accuracy and simulated trading profitability. Our combination of adaptive segmentation and pattern-conditioned forecasting enables more robust and responsive cryptocurrency price prediction.
The impact of sentiment analysis of comments on social networks such as X (Twitter) on the cryptocurrency market’s behavior has been proven. Also, traditional sentiment analysis and not considering the possible aspects of tweets can cause the deep model to be misleading in predicting the price trend of cryptocurrencies. In this research, a model using transfer learning and the combination of pretrained DistilBERT networks, BiGRU deep neural network, and attention layer is presented to analyze the sentiments based on the aspect of tweets and predict the price trend of eight cryptocurrencies. These tweets are the opinions of 70 cryptocurrency expert influencers. After preprocessing, these tweets are injected into the hybrid model of DistilBERT, BiGRU, and attention layer (HDBA) to extract the aspect and determine the polarity of each aspect. The output of the HDBA model is entered into the combined model of BiGRU and the attention layer (HBA) to predict the price trend of each cryptocurrency in intervals of 1–10 days. The output of the HBA model is the best time interval of the influence of the sentiments of tweets on the price trend of cryptocurrencies. The results show that the HDBA model has improved the performance of the aspect‐based sentiment analysis task by an average of 3% in the benchmark datasets. The results of the HBA model also show that this model has been able to predict the best time frame of the impact of sentiments on the behavior of the cryptocurrency market with an average accuracy of 68% and a precision of 73%.
In today's society, news and advertisements have a special place in the growth and development of society. By specifying the main words of the ad, you can understand its general meaning. Preparing these words in the traditional way requires time and specialized knowledge about the subject of the text. Ideakav site is a system that collects Telegram messages and advertisements. The requirement of the idea search system was to extract keywords from the advertisements published in Telegram. The quality of extracted keywords plays a significant role in improving SEO and advertising statistics. By using embedding algorithms, it is possible to extract colloquial conversations and the semantic structure of the text, therefore, it is useful in identifying keywords in Telegram ads that are often published in popular form. In this research, a model of word embedding has been implemented using the data of the idea mining system. The innovation used in this research is created by combining word embedding methods, word frequency and word position. The embedding model is created from two-word words. Creating a model of two-word words is because most of the keywords consist of two words or more. In order to better display the evaluations, the IK model (proposed model) has been compared with statistical methods and graph-based methods, and the obtained results show that the combination of the two-gram IK model has produced a better performance in extracting keywords than other methods.
Organizing and managing cryptocurrency portfolios and decision-making on transactions is crucial in this market. Optimal selection of assets is one of the main challenges that requires accurate prediction of the price of cryptocurrencies. In this work, we categorize the financial time series into several similar subseries to increase prediction accuracy by learning each subseries category with similar behavior. For each category of the subseries, we create a deep learning model based on the attention mechanism to predict the next step of each subseries. Due to the limited amount of cryptocurrency data for training models, if the number of categories increases, the amount of training data for each model will decrease, and some complex models will not be trained well due to the large number of parameters. To overcome this challenge, we propose to combine the time series data of other cryptocurrencies to increase the amount of data for each category, hence increasing the accuracy of the models corresponding to each category.
We present a neural model for sentiment analysis of social network texts with a special focus on cryptocurrency-related content using deep transfer learning. A challenge of deep learning is its need for abundant data. Therefore, we use relation-based transfer learning to analyze low-volume sentiment on SemEval public data. We use the pre-trained BERT-Base model to extract features from the source domain dataset, which is very similar to the extracted tweets of cryptocurrencies. We extract the tweets of expert influencers in cryptocurrency, referred to as the target domain. The extracted features are then injected into an implicit neural network in the target domain. The implicit neural network (INN) was recently designed to work on continuous data such as photos and videos and has not yet been used on text data. Our results show that using implicit neural networks in the text has synergistic effects due to its ability to process continuous and intermittent data. In addition, using rich data in the source domain has caused our proposed model to achieve accuracy of 88% and a loss rate of about 3% on the SemEval data, which is 3% improvement compared to the state-of-the-art. We use the time decay method to slowly change the learning rate in the Adaptive Moment Estimation weighted optimization algorithm. In our proposed model, this technique has helped to reduce the number of neural network cycles by at least 10% to reach convergence.
Objectives Due to the limitations of Twitter, the expansion of Telegram channels, and the Telegram API's easy use, Telegram comments have become prevalent. Telegram is one of the most popular social networks, unlike Twitter, which has no restrictions on sending messages, and experts can share their opinions and media. Some of these channels, managed by influencers of large companies, are very influential in the behavior of the market on various stocks, including cryptocurrencies. In this research, the opinion collection of 10 famous Telegram channels regarding the analysis of cryptocurrencies has been extracted. The sentiments of these opinions have been analyzed using the HDRB model. HDRB is a hybrid model of RoBERTa deep neural network, BiGRU, and attention layer used for sentiment analysis (SA). Analyzing the sentiments of these opinions is very important for understanding the future behavior of the market and managing the stock portfolio. The opinions of this dataset, published by experts in the field of cryptocurrencies, are precious, unlike the opinions that are extracted only by using the hashtag of the names of cryptocurrencies. On the other hand, the dataset related to cryptocurrencies, which has the opinions of experts and the polarity of their feelings, is very rare.Data description The dataset of this research is the sentiments of more than ten popular Telegram channels regarding a wide range of cryptocurrencies. These comments were collected through the Telegram API from December 2023 to March 2024. This data set contains an Excel file containing the text of the comments, the date of comment creation, the number of views, the compound score, the sentiment score, and the type of sentiment polarity. These opinions cover influencer analysis on a wide range of cryptocurrencies. Also, two Word files, one containing the description of the dataset columns and the other Python code for extracting comments from Telegram channels, are included in this dataset.
Stock price movement prediction is challenging due to unpredictable fluctuations and the significant impact of market sentiment and news. Accurate prediction models can enhance investor decision-making and control over stock price movements. Creating a model for predicting high-accuracy stock price movements can improve investor control over stock prices. In this study, a wide range of technical indicators and various aspects of sentiment analysis in tweets were used to predict stock price movement. The impact of the maximum number of positive comments on Tesla stocks on price increases is investigated. Also we proposed a method for adding sentiments to each tweet. Extracted advanced sentiment analysis features such as the number of positive comments, the number of negative comments, the average score of positive comments, the average score of negative comments, daily tweet volume, ratio of positive to negative tweets. Effect of time windows with variate size is investigate. A CNN-LSTM deep neural network is used to predict stock price movement and compared with LSTM and GRU models. According to the results, the proposed CNN-LSTM deep neural network has the best results to predict stock price movement over the 30-day interval.
This article proposes a novel hybrid network integrating three distinct architectures -CNN, GRU, and LSTM- to predict stock price movements. Here with Combining Feature Extraction and Sequence Learning and Complementary Strengths can Improved Predictive Performance. CNNs can effectively identify short-term dependencies and relevant features in time series, such as trends or spikes in stock prices. GRUs designed to handle sequential data. They are particularly useful for capturing dependencies over time while being computationally less expensive than LSTMs. In the hybrid model, GRUs help maintain relevant historical information in the sequence without suffering from vanishing gradient problems, making them more efficient for long sequences. LSTMs excel at learning long-term dependencies in sequential data, thanks to their memory cell structure. By retaining information over longer periods, LSTMs in the hybrid model ensure that important trends over time are not lost, providing a deeper understanding of the time series data. The novelty of the 1D-CNN-GRU-LSTM hybrid model lies in its ability to simultaneously capture short-term patterns and long-term dependencies in time series data, offering a more nuanced and accurate prediction of stock prices. The data set comprises technical indicators, sentiment analysis, and various aspects derived from pertinent tweets. Stock price movement is categorized into three categories: Rise, Fall, and Stable. Evaluation of this model on five years of transaction data demonstrates its capability to forecast stock price movements with an accuracy of 0.93717. The improvement of proposed hybrid model for stock movement prediction over existing models is 12% for accuracy and F1-score metrics.
Software Cost Estimation (SCE) is one of the most widely used and effective activities in project management. In machine learning methods, some features have adverse effects on accuracy. Thus, preprocessing methods based on reducing non-effective features can improve accuracy in these methods. In clustering techniques, samples are categorized into different clusters according to their semantic similarity. Accordingly, in the proposed study, to improve SCE accuracy, first samples are clustered based on original features. Then, a feature selection (FS) technique is separately done for each cluster. The proposed FS method is based on a combination of filter and wrapper FS methods. The proposed method uses both filter and wrapper advantages in selecting effective features of each cluster, with less computational complexity and more accuracy. Furthermore, as the assessment criteria have significant impacts on wrapper methods, a fused criterion has also been used. The proposed method was applied to Desharnais, COCOMO81, COCONASA93, Kemerer, and Albrecht datasets, and the obtained Mean Magnitude of Relative Error (MMRE) for these datasets were 0.2173, 0.6489, 0.3129, 0.4898 and 0.4245, respectively. These results were compared with previous studies and showed improvement in the error rate of SCE.
With the expansion of social networks, sentiment analysis has become one of the hot topics in machine learning. However, in traditional sentiment analysis, the text is considered of a general nature and ignores the different aspects that may exist in the text. This paper presents a hybrid model of transfer deep learning methods for the aspect-oriented sentiment analysis of influencers’ tweets to predict the trend of cryptocurrencies. In the first model, different aspects of tweets are extracted using the Concept Latent Dirichlet Allocation (Concept-LDA). Then, by using the pre-trained RoBERTa network and combining it with the Bidirectional Gated Recurrent Unit (BiGRU) deep learning network and attention layer, sentiments of different aspects of tweets are determined. In the following, the price trend of seven cryptocurrencies, Bitcoin, Ethereum, Binance, Ripple, Dogecoin, Cardano, and Solana, is determined using the historical price and the polarity of tweets with BiGRU combined deep neural network and the attention layer. Also, we used the gridsearch method to select dropout hyper-parameters, learning rate, and the number of GRU units, and the Akaike Information Criterion (AIC) criterion confirmed the results of this proposed combination. The results show that the proposed model in the aspect-based sentiment analysis section has been able to achieve 5.94% accuracy and 9.9% improvement in the f1-score on the SemEval 2015 dataset and 2.61% improvement on the SemEval 2016 dataset in f1-score compared to the state-of-arts. Also, the results of predicting the price trend of cryptocurrencies show that the proposed model has correctly recognized the price trend in the next five days in 77% of cases according to the ROC-AUC criterion.
In the wake of the novel Covid-19 disease pandemic, the global economy has been affected and health crises are widespread. The disease is still incurable, and no effective treatment exists for it. During the Covid-19 crisis, drug repurposing has proven to be an effective treatment strategy. Drug repositioning is an approach to finding effective drugs for treating new diseases by discovering new efficacy of existing drugs. Studies on drug-virus association can reveal new efficacy. The antiviral drug repositioning problem is defined here as a matrix completion problem in which antiviral drugs go down the rows, while viruses go down the columns. We propose a hybrid model called AutoMF that identifies new drug–virus association. This new hybrid model aims to develop a matrix factorization model with deep learning and use it to predict drug-virus associations for repositioning drugs. In our model, a heterogeneous drug–virus network is used, which combines drug–virus associations, drug-drug similarity matrix, and virus-virus similarity matrix. Matrix factorization can extract latent factors from the drug-virus associations. The sparse nature of the associations may prevent the latent factors from being very effective. To solve this problem, deep learning is used to learn effective latent representations from similarity matrices to incorporate with MF to enhance their latent factor priors. Our model outperforms several recent approaches in comparison to benchmarking tests performed on the DVA dataset. In our approach, we identify antiviral drugs currently being tested in clinical trials or those used currently to treat patients.
Drug–target interaction is crucial in the discovery of new drugs. Computational methods can be used to identify new drug–target interactions at low costs and with reasonable accuracy. Recent studies pay more attention to machine-learning methods, ranging from matrix factorization to deep learning, in the DTI prediction. Since the interaction matrix is often extremely sparse, DTI prediction performance is significantly decreased with matrix factorization-based methods. Therefore, some matrix factorization methods utilize side information to address both the sparsity issue of the interaction matrix and the cold-start issue. By combining matrix factorization and autoencoders, we propose a hybrid DTI prediction model that simultaneously learn the hidden factors of drugs and targets from their side information and interaction matrix. The proposed method is composed of two steps: the pre-processing of the interaction matrix, and the hybrid model. We leverage the similarity matrices of both drugs and targets to address the sparsity problem of the interaction matrix. The comparison of our approach against other algorithms on the same reference datasets has shown good results regarding area under receiver operating characteristic curve and the area under precision–recall curve. More specifically, experimental results achieve high accuracy on golden standard datasets (e.g., Nuclear Receptors, GPCRs, Ion Channels, and Enzymes) when performed with five repetitions of tenfold cross-validation. Display graphical of the hybrid model of Matrix Factorization with Denoising Autoencoders with the help side information of drugs and targets for Prediction of Drug-Target Interactions
Finding groups is very useful for marketers to grow their business. Also, that can solve the problem of confusing people in finding groups that are relevant to the user’s interests. Telegram currently has about 500 million monthly active users. It is also known as a social network in many countries such as Iran but does not have the important capabilities of social networks like item recommendation. A content-based recommendation is one of the most successful recommendation techniques based on content correlation. In this paper, we present a new combination model of content based recommendation and clustering method, which mainly consists of three parts. In the first step, the groups and user profiles are extracted from the data using TF-IDF. Second, by clustering groups’ profiles, this model speeds up the identification of groups associated with a user’s profile and assigns a cluster number to them. Finally, to find the most relevant groups, the similarity between the user profile model and the profile of groups with the same cluster is calculated. The role of the proposed algorithm is to improve the performance of the content-based recommendation using the clustering optimization method. The actual dataset of this study includes more than 700,000 supergroups and 70 million users on Telegram. To evaluate the accuracy and performance of the proposed method, we will use precision, recall, and f-measure. Experiments on the proposed method show the effectiveness of recommendations, high accuracy, and speed improvements.
A General Investigation on the Combination of Local and Global Feature Selection Methods for Request Identification on Telegram
Telegram currently has nearly 500 million monthly active users. It is known as a social network in many countries, such as Iran. Due to the diversity and increase of groups in this messenger, it has become difficult to find groups that are related to the users’ interests; so it is better for users to find suitable groups with the help of recommender systems. In this paper, we present a new collaborative filtering method for a group recommender system. It is based on the users’ interests, using and analyzing the graph of users’ membership. The proposed method mainly consists of two parts. In the first step, the main user and top similar users are extracted, then we list the top similar users’ groups. For each similar user, we calculate a score based on the public and the number of user’s groups for any group in the list. We consider the total score of all the top similar users for each group in the list as the group rank. In the second step, due to the existence of groups with only one similar user, the score of popular groups changes according to the number of dissimilar members in each group. Experiments have been performed to determine the effectiveness of these methods on Telegram data. The actual data of this research includes more than 700,000 supergroups and 70 million Telegram users. To evaluate the accuracy and performance of this research precision, recall and f-measure are used and the results confirm the usefulness and effectiveness of the proposed method.
Software Development Effort Estimation (SDEE) can be interpreted as a set of efforts to produce a new software system. To increase the estimation accuracy, the researchers tried to provide various machine learning regressors for SDEE. Kernel Ridge Regression (KRR) has demonstrated good potentials to solve regression problems as a powerful machine learning technique. Gravitational Search Algorithm (GSA) is a metaheuristic method that seeks to find the optimal solution in complex optimization problems among a population of solutions. In this article, a hybrid GSA algorithm is presented that combines Binary-valued GSA (BGSA) and the real-valued GSA (RGSA) in order to optimize the KRR parameters and select the appropriate subset of features to enhance the estimation accuracy of SDEE. Two benchmark datasets are considered in the software projects domain for assessing the performance of the proposed method and similar methods in the literature. The experimental results on Desharnais and Albrecht datasets have confirmed that the proposed method significantly increases the accuracy of the estimation comparing some recently published methods in the literature of SDEE.
Today, recommender systems are used in many different businesses to find items of interest to users. The use of these systems is widely found in online economic systems and social networks. Therefore, using these systems in the messaging environment will cause changes and transformations for marketing. Telegram is a cloud-based messenger with more than 500 million monthly active users. This messenger has a relatively acceptable position compared to its other competitors, because the security and features provided in it, have made it different from other messengers and close to social networks. One of the most popular features of messengers is groups. Many marketers are looking for groups that fit their field. One of the main gaps in the messengers regarding advertising and marketing to expand businesses is the impossibility of finding social groups. In this paper, a new method for group recommendation in the Telegram is presented. This method, by receiving a set of users, analyzes their groups and recommends a list of ranked groups. The proposed method is created by combining the previous two methods in the field of group recommendation and computational modeling of numerical variables obtained from each group. This study is dependent on the information of all users due to the use of the membership graph, and the behavior of the system changes by the information extracted from the users. The results of experimental experiments show a significant reduction of RMSE and MAE in the proposed method compared to the previous two methods.
Telegram is a cloud-based instant messenger with more than 500 million monthly active users. This messenger is very popular among Iranians, as more than 50 million Telegram users are Iranians. Telegram is used as a social network in Iran because it offers features beyond a simple messenger, but does not offer all the features of social networks, including user recommendation. In this paper, investigating a real dataset crawled from Telegram, we have provided a hybrid method using the user membership graph and group characteristics to recommend the user in Telegram. The membership graph connects users based on membership in the same groups. Also, the characteristics for each group are indicated by the name and description of that group in Telegram. We created a bag of words for each group using natural language processing methods, then combined the bag of words for each group with the results of the membership graph processing. Finally, users are recommended based on the list of groups obtained by the combination. The data used in this paper include more than 900,000 groups and 120 million users. Evaluation of the proposed method separately on two categories of Telegram specialized groups shows the model integration and error reduction for the first category to 0.009 and the second category to 0.016 in RMSE.