This study introduces a novel evaluation framework for predicting web page performance, utilizing state-of-the-art machine learning algorithms to enhance the accuracy and efficiency of web quality assessment. We systematically identify and analyze 59 key attributes that influence website performance, derived from an extensive literature review spanning from 2010 to 2024. By integrating a comprehensive set of performance metrics—encompassing usability, accessibility, content relevance, visual appeal, and technical performance—our framework transcends traditional methods that often rely on limited indicators. Employing various classification algorithms, including Support Vector Machines (SVMs), Logistic Regression, and Random Forest, we compare their effectiveness on both original and feature-selected datasets. Our findings reveal that SVMs achieved the highest predictive accuracy of 89% with feature selection, compared to 87% without feature selection. Similarly, Random Forest models showed a slight improvement, reaching 81% with feature selection versus 80% without. The application of feature selection techniques significantly enhances model performance, demonstrating the importance of focusing on impactful predictors. This research addresses critical gaps in the existing literature by proposing a methodology that utilizes newly extracted features, making it adaptable for evaluating the performance of various website types. The integration of automated tools for evaluation and predictive capabilities allows for proactive identification of potential performance issues, facilitating informed decision-making during the design and development phases. By bridging the gap between predictive modeling and optimization, this study contributes valuable insights to practitioners and researchers alike, establishing new benchmarks for future investigations in web page performance evaluation.
The detection of spam reviews in multilingual environments remains a challenging task due to linguistic diversity, data imbalance, and semantic complexity. This paper proposes a novel hybrid model that integrates Twin Support Vector Machine (TwinSVM) with Harris Hawks Optimization (HHO) for simultaneous parameter optimization and feature selection. To enhance semantic understanding, sentiment-based features are incorporated alongside pre-trained word embedding models-BERT, FastText, and MUSE-across English, Arabic, and Spanish datasets. Our approach generates 24 high-quality datasets using embeddings with 100 and 400 dimensions, including a combined multilingual set. Experimental results demonstrate that our proposed HHO-TwinSVM model consistently outperforms conventional classifiers and metaheuristic-enhanced SVMs, achieving accuracy improvements of up to 9.44% and enhanced robustness in low-resource languages. This integrated framework represents a scalable and adaptable solution for multilingual spam detection. Four detailed experiments were conducted in this study, each designed to address and demonstrate a specific aspect of the proposed approach. Across all experiments, the method outperformed existing algorithms, achieving impressive accuracy rates of 92.9741 %, 89.0314%, 80.3580%, and 85.0859% on Arabic, English, Spanish, and multilingual datasets, respectively. Subsequently, sentiment analysis features were incorporated to further enhance detection performance, resulting in improvements of 1.0994%, 2.6674%, 9.4430%, and 8.7448%, respectively. A comprehensive analysis of the experimental results, including the influence of reviews and sentiment features, is also presented.
The global community is awaiting the advent of a self-driving vehicle that is safe, reliable, and capable of navigating a diverse range of road conditions and terrains. This requires a lot of research, study, and optimization. Thus, this work focused on implementing, training, and optimizing a convolutional neural network (CNN) model, aiming to predict the steering angle during driving (one of the main issues). The considered dataset comprises images collected inside a car-driving simulator and further processed for augmentation and removal of unimportant details. In addition, an innovative data-balancing process was previously performed. A CNN model was trained with the dataset, conducting a comparison between several different standard optimizers. Moreover, evolutionary optimization was applied to optimize the model’s weights as well as the optimizers themselves. Several experiments were performed considering different approaches of genetic algorithms (GAs) along with other optimizers from the state of the art. The obtained results demonstrate that the GA is an effective optimization tool for this problem.
Due to the escalating network throughput and security risks, the exploration of intrusion detection systems (IDSs) has garnered significant attention within the computer science field. The majority of modern IDSs are constructed using deep learning techniques. Nevertheless, these IDSs still have shortcomings where most datasets used for IDS lies in their high imbalance, where the volume of samples representing normal traffic significantly outweighs those representing attack traffic. This imbalance issue restricts the performance of deep learning classifiers for minority classes, as it can bias the classifier in favor of the majority class. To address this challenge, many solutions are proposed in the literature. TDCGAN is an innovative Generative Adversarial Network (GAN) based on a model-driven approach used to address imbalanced data in the IDS dataset. This paper investigates the performance of TDCGAN by employing it to balance data across four benchmark IDS datasets which are CIC-IDS2017, CSE-CIC-IDS2018, KDD-cup 99, and BOT-IOT. Next, four machine learning methods are employed to classify the data, both on the imbalanced dataset and on the balanced dataset. A comparison is then conducted between the results obtained from each to identify the impact of having an imbalanced dataset on classification accuracy. The results demonstrated a notable enhancement in the classification accuracy for each classifier after the implementation of the TDCGAN model for data balancing.
Platelet-derived growth factor-D (PDGF-D) is abundantly expressed in ocular diseases. Yet, it remains unknown whether and how PDGF-D affects ocular cells or cell-cell interactions in the eye. In this study, using single-cell RNA sequencing (scRNA-seq) and a mouse model of PDGF-D overexpression in retinal pigment epithelial (RPE) cells, we found that PDGF-D overexpression markedly upregulated the key immunoproteasome genes, leading to increased antigen processing/presentation capacity of RPE cells. Also, more than 6.5-fold ligand-receptor pairs were found in the PDGF-D overexpressing RPE-choroid tissues, suggesting markedly increased cell-cell interactions. Moreover, in the PDGF-D-overexpressing tissues, a unique cell population with a transcriptomic profile of both stromal cells and antigen-presenting RPE cells was detected, suggesting PDGF-D-induced epithelial-mesenchymal transition of RPE cells. Importantly, administration of ONX-0914, an immunoproteasome inhibitor, suppressed choroidal neovascularization (CNV) in a mouse CNV model in vivo. Together, we show that overexpression of PDGF-D increased pro-angiogenic immunoproteasome activities, and inhibiting immunoproteasome pathway may have therapeutic value for the treatment of neovascular diseases.
Messaging platforms are applications, generally mediated by an app, desktop program or the web, mainly used for synchronous communication among users. As such, they have been widely adopted officially by higher education establishments, after little or no study of their impact and perception by the teachers. We think that the introduction of these new tools and the opportunities and challenges they have needs to be studied carefully in order to adopt the model, as well as the tool, that is the most adequate for all parties involved. We already studied the perception of these tools by students, in this paper we examine the teachers' experiences and perceptions through a survey that we validated with peers, and what they think these tools should make or serve so that it enhances students learning and helps them achieve their learning objectives. The survey has been distributed among tertiary education teachers, both in universitary and other kind of tertiary establishments, based in Spain (mainly) and Spanish-speaking countries. We have focused on collecting teachers' preferences and opinions on the introduction of messaging platforms in their day-to-day work, as well as other services attached to them, such as chatbots. What we intend with this survey is to understand their needs and to gather information about the various educational use cases where these tools could be valuable. In addition, an analysis of how and when teachers' opinions towards the use of these tools varies across gender, experience, and their discipline of specialization is presented. The key findings of this study highlight the factors that can contribute to the advancement of the adoption of messaging platforms and chatbots in higher education institutions to achieve the desired learning outcomes.
Computational prediction of cell-cell interactions (CCIs) is becoming increasingly important for understanding disease development and progression. We present a benchmark study of available CCI prediction tools based on single-cell RNA sequencing (scRNA-seq) data. By comparing prediction outputs with a manually curated gold standard for idiopathic pulmonary fibrosis (IPF), we evaluated prediction performance and processing time of several CCI prediction tools, including CCInx, CellChat, CellPhoneDB, iTALK, NATMI, scMLnet, SingleCellSignalR, and an ensemble of tools. According to our results, CellPhoneDB and NATMI are the best performer CCI prediction tools, among the ones analyzed, when we define a CCI as a source-target-ligand-receptor tetrad. In addition, we recommend specific tools according to different types of research projects and discuss the possible future paths in the field.
Intrusion Detection Systems (IDSs) are a primary research area in Cybersecurity nowadays. These are programs or methods designed to monitor and analyze network traffic aiming to identify suspicious patterns/attacks. MSNM (Multivariate Statistical Network Monitoring) is a state-of-the-art algorithm capable of detecting various security threats in real network traffic data with high performance. However, semi-supervised MSNM heavily relies on a set of weights, whose values are usually determined using a relatively simple optimization algorithm. This work proposes the application of various Evolutionary Algorithm approaches to optimize this set of variables and improve the performance of MSNM against four types of attacks using the UGR’16 dataset (includes real network traffic flows). Furthermore, we analyzed the performance of a Particle Swarm Optimization approach and a Simulated Annealing algorithm, as a baseline. The results obtained are very promising and show that EAs are a great tool for enhancing the performance of this IDS.
Online reviews are important information that customers seek when deciding to buy products or services. Also, organizations benefit from these reviews as essential feedback for their products or services. Such information required reliability, especially during the Covid-19 pandemic which showed a massive increase in online reviews due to quarantine and sitting at home. Not only the number of reviews was boosted but also the context and preferences during the pandemic. Therefore, spam reviewers reflect on these changes and improve their deception technique. Spam reviews usually consist of misleading, fake, or fraudulent reviews that tend to deceive customers for the purpose of making money or causing harm to other competitors. Hence, this work presents a Weighted Support Vector Machine (WSVM) and Harris Hawks Optimization (HHO) for spam review detection. The HHO works as an algorithm for optimizing hyperparameters and feature weighting. Three different language corpora have been used as datasets, namely English, Spanish, and Arabic in order to solve the multilingual problem in spam reviews. Moreover, pre-trained word embedding (BERT) has been applied alongside three-word representation methods (NGram-3, TFIDF, and One-hot encoding). Four experiments have been conducted, each focused on solving and demonstrating different aspects. In all experiments, the proposed approach showed excellent results compared with other state-of-the-art algorithms. In other words, the WSVM-HHO achieved an accuracy of 88.163%, 71.913%, 89.565%, and 84.270%, for English, Spanish, Arabic, and Multilingual datasets, respectively. Further, a deep analysis has been conducted to investigate the context of reviews before and after the COVID-19 situation. In addition, it has been generated to create a new dataset with statistical features and merge its previous textual features for improving detection performance.
This paper presents a study on the creation of a tool to help powerlifting athletes and coaches, as well as bodybuilders and other amateur gym athletes, to analyse their data and obtain useful information regarding the athlete’s performance. The tool should also predict future personal records in lifting for both raw (non-equipped) and non-raw (equipped) attempts, and their various exercises. In order to achieve this, a dataset with entries of around 500 k lifters and more than 20 k official powerlifting competitions was used. Among those entries, biometric variables of the lifters and the weights they lift in each of the three movements of this sport discipline were included: squat, bench press, and deadlift. We applied data preprocessing and visualising as well as data splitting and scaling techniques in order to train the machine learning models that are used to make the predictions. Lastly, the best predictive models were used in the implemented tool.
An intrusion detection system (IDS) plays a critical role in maintaining network security by continuously monitoring network traffic and host systems to detect any potential security breaches or suspicious activities. With the recent surge in cyberattacks, there is a growing need for automated and intelligent IDSs. Many of these systems are designed to learn the normal patterns of network traffic, enabling them to identify any deviations from the norm, which can be indicative of anomalous or malicious behavior. Machine learning methods have proven to be effective in detecting malicious payloads in network traffic. However, the increasing volume of data generated by IDSs poses significant security risks and emphasizes the need for stronger network security measures. The performance of traditional machine learning methods heavily relies on the dataset and its balanced distribution. Unfortunately, many IDS datasets suffer from imbalanced class distributions, which hampers the effectiveness of machine learning techniques and leads to missed detection and false alarms in conventional IDSs. To address this challenge, this paper proposes a novel model-based generative adversarial network (GAN) called TDCGAN, which aims to improve the detection rate of the minority class in imbalanced datasets while maintaining efficiency. The TDCGAN model comprises a generator and three discriminators, with an election layer incorporated at the end of the architecture. This allows for the selection of the optimal outcome from the discriminators' outputs. The UGR'16 dataset is employed for evaluation and benchmarking purposes. Various machine learning algorithms are used for comparison to demonstrate the efficacy of the proposed TDCGAN model. Experimental results reveal that TDCGAN offers an effective solution for addressing imbalanced intrusion detection and outperforms other traditionally used oversampling techniques. By leveraging the power of GANs and incorporating an election layer, TDCGAN demonstrates superior performance in detecting security threats in imbalanced IDS datasets.
Although VEGF-B was discovered as a VEGF-A homolog a long time ago, the angiogenic effect of VEGF-B remains poorly understood with limited and diverse findings from different groups. Notwithstanding, drugs that inhibit VEGF-B together with other VEGF family members are being used to treat patients with various neovascular diseases. It is therefore critical to have a better understanding of the angiogenic effect of VEGF-B and the underlying mechanisms. Using comprehensive in vitro and in vivo methods and models, we reveal here for the first time an unexpected and surprising function of VEGF-B as an endogenous inhibitor of angiogenesis by inhibiting the FGF2/FGFR1 pathway when the latter is abundantly expressed. Mechanistically, we unveil that VEGF-B binds to FGFR1, induces FGFR1/VEGFR1 complex formation, and suppresses FGF2-induced Erk activation, and inhibits FGF2-driven angiogenesis and tumor growth. Our work uncovers a previously unrecognized novel function of VEGF-B in tethering the FGF2/FGFR1 pathway. Given the anti-angiogenic nature of VEGF-B under conditions of high FGF2/FGFR1 levels, caution is warranted when modulating VEGF-B activity to treat neovascular diseases.
This paper presents a preliminary study on the application of evolutionary algorithms for the optimisation of the behaviour of autonomous agents for 1 vs 1 fighting games. Different evolutionary schemes have been proposed and quite promising results have been obtained, which leave room for obtaining very competitive agents following this methodology.
This review discusses our current understanding of chromatin biology and bioinformatics under the unifying concept of “chromatin hubs.” The first part reviews the biology of chromatin hubs, including chromatin–chromatin interaction hubs, chromatin hubs at the nuclear periphery, hubs around macromolecules such as RNA polymerase or lncRNAs, and hubs around nuclear bodies such as the nucleolus or nuclear speckles. The second part reviews existing computational methods, including enhancer–promoter interaction prediction, network analysis, chromatin domain callers, transcription factory predictors, and multi-way interaction analysis. We introduce an integrated model that makes sense of the existing evidence. Understanding chromatin hubs may allow us (i) to explain long-unsolved biological questions such as interaction specificity and redundancy of mechanisms, (ii) to develop more realistic kinetic and functional predictions, and (iii) to explain the etiology of genomic disease.
Gene Set Analysis (GSA) is one of the most commonly used strategies to analyze omics data. Hundreds of GSA-related papers have been published, giving birth to a GSA field in Bioinformatics studies. However, as the field grows, it is becoming more difficult to obtain a clear view of all available methods, resources, and their quality. In this paper, we introduce a web platform called “GSA Central” which, as its name indicates, acts as a focal point to centralize GSA information and tools useful to beginners, average users, and experts in the GSA field. “GSA Central” contains five different resources: A Galaxy instance containing GSA tools (“Galaxy-GSA”), a portal to educational material (“GSA Classroom”), a comprehensive database of articles (“GSARefDB”), a set of benchmarking tools (“GSA BenchmarKING”), and a blog (“GSA Blog”). We expect that “GSA Central” will become a useful resource for users looking for introductory learning, state-of-the-art updates, method/tool selection guidelines and insights, tool usage, tool integration under a Galaxy environment, tool design, and tool validation/benchmarking. Moreover, we expect this kind of platform to become an example of a “thematic platform” containing all the resources that people in the field might need, an approach that could be extended to other bioinformatics topics or scientific fields.
Digital Collectible Cards Games such as Hearthstone have become a very prolific test-bed for Artificial Intelli-gence algorithms. The main researches have focused on the implementation of autonomous agents (bots) able to effectively play the game. However, this environment is also very attractive for the use of Data Mining (DM) and Machine Learning (ML) techniques, for analysing and extracting useful knowledge from game data. The objective of this work is to apply existing Game Mining techniques in order to study more than 600,000 real decks (groups of cards) created by players with many different skill levels. Data visualisation and analysis tools have been applied, namely, Graph representations and Clustering techniques. Then, an expert player has conducted a deep analysis of the results yielded by these methods, aiming to identify the use of standard -and well-known - ar-chetypes defined by the play methods will also make it possible for the expert to discover hidden relationships between cards that could lead to finding better combinations of them, enhancing players' decks or, otherwise, identify unbalanced cards that could lead to a disappointing game experience. Moreover, although this work is mostly focused on data analysis and visualization, the obtained results can be applied to improve Hearthstone Bots' behaviour, e.g. predicting opponent's actions after identifying a specific archetype in his/her deck.
During the recent COVID-19 pandemic, people were forced to stay at home to protect their own and others’ lives. As a result, remote technology is being considered more in all aspects of life. One important example of this is online reviews, where the number of reviews increased promptly in the last two years according to Statista and Rize reports. People started to depend more on these reviews as a result of the mandatory physical distance employed in all countries. With no one speaking to about products and services feedback. Reading and posting online reviews becomes an important part of discussion and decision-making, especially for individuals and organizations. However, the growth of online reviews usage also provoked an increase in spam reviews. Spam reviews can be identified as fraud, malicious and fake reviews written for the purpose of profit or publicity. A number of spam detection methods have been proposed to solve this problem. As part of this study, we outline the concepts and detection methods of spam reviews, along with their implications in the environment of online reviews. The study addresses all the spam reviews detection studies for the years 2020 and 2021. In other words, we analyze and examine all works presented during the COVID-19 situation. Then, highlight the differences between the works before and after the pandemic in terms of reviews behavior and research findings. Furthermore, nine different detection approaches have been classified in order to investigate their specific advantages, limitations, and ways to improve their performance. Additionally, a literature analysis, discussion, and future directions were also presented.
Real-Time Strategy (RTS) games are well-known for their substantially large combinatorial decision and state spaces, responsible for creating significant challenges for search and machine learning techniques. Exploiting domain knowledge to assist in navigating the expansive decision and state spaces could facilitate the emergence of competitive RTS game-playing agents. Usually, domain knowledge can take the form of expert traces or expert-authored scripts. A script encodes a strategy conceived by a human expert and can be used to steer a search algorithm, such as Monte Carlo Tree Search (MCTS), towards high-value states. However, a script is coarse by nature, meaning that it could be subject to exploitation and poor low-level tactical performance. We propose to perceive scripts as a collection of heuristics that can be parameterized and combined to form a wide array of strategies. The parameterized heuristics mold and filter the decision space in favor of a strategy expressed in terms of parameters. The proposed agent, ParaMCTS, implements several common heuristics and uses NaiveMCTS to search the downsized decision space; however, it requires a preceding manual parameterization step. A genetic algorithm is proposed for use in an optimization phase that aims to replace manual tuning and find an optimal set of parameters for use by EvoPMCTS, the evolutionary counterpart of ParaMCTS. Experimentation results using the mu RTS testbed show that EvoPMCTS outperforms several state-of-the-art agents across multiple maps of distinct layouts.
The critical factors regulating stem cell endothelial commitment and renewal remain not well understood. Here, using loss- and gain-of-function assays together with bioinformatic analysis and multiple model systems, we show that PDGFD is an essential factor that switches on endothelial commitment of embryonic stem cells (ESCs). PDGFD genetic deletion or knockdown inhibits ESC differentiation into EC lineage and increases ESC self-renewal, and PDGFD overexpression activates ESC differentiation towards ECs. RNA sequencing reveals a critical requirement of PDGFD for the expression of vascular-differentiation related genes in ESCs. Importantly, PDGFD genetic deletion or knockdown increases ESC self-renewal and decreases blood vessel densities in both embryonic and neonatal mice and in teratomas. Mechanistically, we reveal that PDGFD fulfills this function via the MAPK/ERK pathway. Our findings provide new insight of PDGFD as a novel regulator of ESC fate determination, and suggest therapeutic implications of modulating PDGFD activity in stem cell therapy.
J. Merelo合作论文数Dept. of Computer Technology and Architecture;Universidad de Granada138
A. J. Fernández合作论文数Universidad de Malaga;Departamento de Lenguajes and Ciencias de la Computacion7