
Metabarcoding is a technique for analysing DNA sequences that target specific gene regions and plays a crucial role in the identification and classification of different organisms. In particular, 16S rRNA metabarcoding enables the elucidation of complex bacterial and archaeal communities in food. This research presents a novel dataset obtained by metabarcoding analysis of 16S rRNA aimed at elucidating the microbial dynamics of cooked, ready-to-eat ham products over a defined storage period. At the centre of our investigation is the application of association rule mining, an unsupervised machine learning approach in data mining, to uncover latent patterns and relationships within the dataset. At the taxonomic “family” level, our analysis shows a strong correlation between the presence of Bacillaceae and Staphylococcaceae with a support of 92%. This finding highlights the consistent co-occurrence of these microbial families with a confidence level of 96%, meaning that the presence of Bacillaceae strongly predicts the presence of Staphylococcaceae. Furthermore, at the genus level, a significant relationship is observed between Brochotrix and Arthrobacter, with both genera co-occurring in approximately 85% of samples in the dataset. Notably, the high confidence level of 98% suggests a strong association, suggesting that the presence of Brochotrix reliably predicts the presence of Arthrobacter. These results provide valuable insights into microbial dynamics in food and demonstrate the effectiveness of using advanced data mining techniques in deciphering complex food ecosystems interactions.
In this paper we propose an autoencoder-based End-to-End (E2E) Deep Learning for symbol decoding in fiber optical communications systems. The autoencoder was trained with data encoded in two different ways. First, one-hot encoding which corresponds to the most common approach to train the network in order to optimize the decision regions of each symbol. Then, we introduce the binary coding where the bit sequences that represent the symbols, are positioned in such a way to minimize the error and compensate the distortions introduced by the optical channel. The focus of the study are the generated constellations, symbol decision regions and symbol bit labeling. The autoencoder contributes to better visualization and analysis of the resulting decision regions for scenarios with high and low nonlinearities. It improves the interpretability of the symbol decoding. To the best of our knowledge, this is the first time when a ML-based approach (i.e the autoencoder) estimates not only the ideal locations of the decoded symbols in various constellations (4, 16, 64) but also optimizes the geometrical decision regions of the symbols in the constellation.
Product Manufacturing Information is the key element in the digital transformation of manufacturing. It facilitates the transition from 2D drawing-based workflows to 3D model-based workflows, unlocking downstream automating tasks in the CAx-process chain. However, the current methods of enhancing 3D-CAD models with PMI, such as manual creation or rule-based approaches, are labor-intensive, error-prone, and inefficient, particularly when applied to numerous similar models. To address these challenges, this work proposes the development of a graph-based deep learning system to automate the PMI creation in 3D-CAD models. A new graph representation for 3D-CAD models incorporating PMI is presented. A data generation pipeline and two homogeneous graph neural network-based methods, node classification and link prediction, are then proposed to predict the existence of PMI for a given 3D-CAD model. Results demonstrate the performance of node classification for predicting PMI on individual faces of unseen 3D-CAD models while link prediction shows the capability of predicting PMI between two faces of a 3D-CAD model. These findings form the groundwork for building a fully automated recommendation system for PMI in real-world industry applications.
This article discusses the basic principles of developing and using mobile agents, which are software components that can move between network nodes to perform various tasks. Describes the architecture of mobile agents, including an agent-based core, migration mechanisms, and runtime, and discusses popular tools and technologies for their development, such as Java Aglets, JADE, Mobile-C, and Voyager. The article analyzes the key benefits of mobile agents, including flexibility, reduced network traffic, and increased reliability, and discusses security and management challenges. In modern pedagogy, there is a trend towards the use of mobile technologies with the introduction of elements of artificial intelligence. Today's mobile learning apps are increasingly using artificial intelligence to personalize the learning experience. However, traditional approaches based on centralized algorithms face difficulties in scaling and adapting to dynamic user needs. Mobile agents offer a new approach to creating more flexible and adaptive learning mobile apps. In this article, we explore the potential of mobile agents in the context of training, looking at their architecture, benefits, and limitations. The use of mobile agents in mobile learning apps represents a new paradigm for personalized learning, aimed at increasing the efficiency of the educational process and improving the user experience. Mobile agents are autonomous software entities that can move between different network nodes and perform tasks tailored to the needs and preferences of users. In learning mobile apps, they can collect data on student behavior and progress, provide personalized recommendations, offer personalized learning plans, and provide real-time, interactive support. In the course of the study, the EduAgent mobile application was developed. The results of the study show that the use of mobile agents contributes to improving the perception of the material, increasing student engagement and motivation. Interactive modules, video lessons, quizzes and assignments, supported by the work of agents, allow students to have a more tailored and interactive educational experience. Progress monitoring with granular reports and progress graphs helps students track their achievements and customize the learning process to their needs. Built-in mobile agent support provides instant assistance and access to training materials. Thus, the integration of mobile agents into educational mobile applications opens up new opportunities for creating more effective and personalized learning environments, which makes this approach a promising direction for further research and development. The results of the study emphasize the prospects of this technology and the need to overcome existing challenges for its successful implementation.
This paper presents the first published application of multiple existing machine learning methods to a subset of features taken from the Profiles of Individual Radicalization in the United States (PIRUS) database to predict the feature ‘violent’. The best-performing model in terms of accuracy is the Hist Gradient Boosting model, with an accuracy of 89.06%, which is an improvement of more than 2.5% compared to the benchmark application. Permutation Feature Importance (PFI) and the explanation framework SHAP were then applied to explain the model predictions. Using both of these techniques together allows for a holistic view of both the model's inner workings and the impact of the features on the results.
Fine particulate matter (PM2.5) air pollution is a global public health crisis, responsible for millions of premature deaths annually. Accurate prediction of PM2.5 concentrations is paramount for timely alerts, effective mitigation, and informed policy decisions. Chronic exposure to high levels of PM2.5 is associated with increased mortality rates from heart and lung diseases. This is especially concerning for vulnerable populations such as children, the elderly, and individuals with pre-existing health conditions. To address this challenge, accurate and timely prediction of PM2.5 levels is essential. Traditional linear models often fall short in capturing the complex interactions between multiple factors influencing air quality. Therefore, this study employs Random Forest Regression, a robust ensemble learning method, enhanced with polynomial features using meteorological data from the U.S. Environmental Protection Agency's (EPA) Air Quality System (AQS) to improve the predictive accuracy. Random Forest Regression is particularly suitable for this task due to its ability to handle high-dimensional data and capture non-linear relationships. The model's ability to capture non-linear relationships and interactions between predictors is thoroughly evaluated. We leverage the extensive AQS dataset, discuss the intricacies of feature engineering, present rigorous model validation, and explore the implications of our findings for air quality management and public health. By incorporating polynomial features, the model aims to better capture the complex interactions among variables such as temperature, humidity, wind speed, and other environmental factors that influence PM2.5 levels. This study not only highlights the effectiveness of using advanced machine learning techniques to enhance air quality predictions but also underscores the critical role of accurate forecasting in protecting public health and guiding policy decisions. Through comprehensive analysis and validation, we demonstrate the potential of this approach to provide more reliable predictions, thus supporting efforts to mitigate the adverse effects of air pollution on a global scale. As air quality continues to be a pressing global issue, innovative approaches such as the one presented in this study are essential for ensuring a healthier and more sustainable future.
In the modern world, any development is impossible without the development of new high-tech automated systems in various subject areas. In this study, three different algorithms were considered that use effective motion control of a wheeled robot. Controlling the movement of a wheeled robot is a complex task that requires an accurate algorithm which allows adaptation to various conditions and tasks. The key is to select an appropriate algorithm that matches the task requirements and the robot's characteristics. Simulators are considered the most effective in the field of modeling robotic systems, complexes and individual objects. When using any simulation, it is important to know the factors that significantly influence the final result. The influence of one of these factors is discussed in this article.
Smart grid monitoring in IoT environments demands robust fault tolerance mechanisms to ensure uninterrupted operation and data accuracy. The integration of advanced machine learning with fault-tolerant strategies in the proposed Intelligent FaultEdge framework represents a significant innovation. Unlike traditional reactive systems, Intelligent FaultEdge adopts a proactive approach, leveraging predictive analytics, anomaly detection, and adaptive reconfiguration to anticipate and manage potential faults in real-time. This proactive framework not only enhances the reliability and resilience of smart grid operations but also minimizes downtime, improves safety, and optimizes energy efficiency. By continuously monitoring grid conditions and dynamically adjusting to evolving scenarios, Intelligent FaultEdge supports sustainable and secure smart grid management. This paper presents a comprehensive framework aligned with the goals of modern intelligent systems, offering an efficient solution for enhancing the reliability and performance of IoT-enabled smart grids. Through its proactive fault management capabilities, Intelligent FaultEdge contributes to advancing smart infrastructure, ensuring dependable energy supply, and supporting the evolution towards more resilient and efficient smart grids.
The study proposes using five benchmark machine-learning models alongside XGBoost, applied for the first time to an existing case study to predict the success of suicide of terrorist attacks. Utilizing data from the Global Terrorism Database (GTD), the study evaluates model effectiveness to aid decision-making for emergency responders and policymakers. Employing explainable Artificial Intelligence (XAI) models like SHAP ensures transparent decision-making processes. XGBoost performed best for accuracy and performance, while LightGBM excelled in explainability, with SHAP providing global and local insights into their decision-making. The primary goal is to enhance user comprehension and facilitate informed decision-making in critical scenarios, prioritizing transparency, and trustworthiness.
Self-paced learning engenders flexibility, individuality and autonomy, allowing learners to acquire knowledge at their convenience. Learners, in fact, increasingly seek flexible and personalized learning experiences. They require systems that not only accommodate but also enhance their learning journeys. This paper presents the combination of artificial intelligence (AI) features and the design of interaction and communication in a learning format that fosters learners' acquisition of comprehensive and sustainable knowledge. The primary focus of our approach is on individuals who are new to the course subject or possess limited prior knowledge. Recognizing the importance of structured and coherent content, our learning format emphasizes uniformity in both appearance and language across text and quiz pages. Such consistency is crucial for building a stable conceptual framework, which facilitates a deeper understanding of the subject matter. By minimizing extraneous cognitive load, we enable learners to concentrate on the intrinsic elements of the content, thereby enhancing their learning experience. A significant aspect of our design is the strategic use of natural language. Consistent language aids in the formation of domain-specific conceptual frameworks, which are essential for sustainable knowledge acquisition.
This paper tackles the challenge of extracting descriptive schemas from JSON data collections, specifically targeting the discovery of tagged unions. Tagged unions are a JSON Schema design pattern where the value of one property dictates conditional subschemas for other sibling properties. By formalizing these implications as conditional functional dependencies, we employ JSON Schema operators such as if-then-else. Our heuristics are designed to avoid overfitting. Promising experiments with our prototype implementation show successful detection of tagged unions in real-world GeoJSON and TopoJSON datasets. Additionally, we explore potential extensions of our approach for future work.
Current logistics companies carry out internal optimizations concerning resource utilization of their freighters. In cooperation with other companies, e.g., to hand over the orders, they use a trustworthy broker that handles this process without exchanging company-related, private information from the logistics companies to not violate confidentiality. Increasing digitalization offers the opportunity to carry out cross-company optimizations and resource allocation. Such optimizations would increase the average load factor of transport providers, and ultimately reduce the overall carbon footprint of the entire transport sector. However, the companies are typically competitors and sharing trade secrets could result in competitors overtaking the market, leading logistics companies to not collaborate. This paper proposes a novel application of swarm intelligence as privacy enabler in the logistics sector. Logistics companies and their resources are modeled as agents in a simulated swarm system where agents interact by following simple local rules. Such agent behavior emerges into a meaningful solution that provides a secure and efficient transportation network. Applying swarm intelligence with local rules from the bottom-up, the paper proposes a secure multi-party optimization in the logistics sector, ensuring that no participant needs to reveal any sensitive business information to any other participant in the system, including any third parties or brokers, thereby immediately overcoming the companies' reluctance to participate. The agent's behavior is inspired by the Artificial Bee Colony Algorithm (ABC). We show that reducing the amount of exchanged data with swarm algorithms still leads to a satisfying optimization in terms of resource utilization of all involved parties.
This paper presents the development and integration of advanced imaging technologies within the ZEMELA platform, a smart agriculture platform tailored for precision viticulture. As the agricultural sector increasingly adopts technological innovations, precision agriculture has emerged as a key player in enhancing crop management and sustainability. The ZEMELA platform utilizes a combination of spectral analysis, thermal imaging, and photogrammetry to provide comprehensive, real-time data on vine health, soil conditions, and crop development. These capabilities enable vineyard managers to make informed decisions that optimize resource use, reduce environmental impact, and increase crop yields. By integrating these imaging technologies into a unified platform, the ZEMELA platform demonstrates the potential for digital technologies to revolutionize traditional farming practices, ensuring sustainability and efficiency. The implementation of this system in Bulgarian vineyards serves as a model for global precision agriculture initiatives, highlighting the crucial role of innovative technologies in the future of farming.
With biometric identification systems becoming increasingly ubiquitous, their complexity is escalating due to the integration of diverse sensors and modalities, aimed at minimizing error rates. The current paradigm for these systems involves hard-coded aggregation instructions, presenting challenges in system maintenance, scalability, and adaptability. These challenges become particularly prominent when deploying new sensors or adjusting security levels to respond to evolving threat models. To address these concerns, this research introduces BioDSSL, a Domain Specific Sensor Language to simplify the integration and dynamic adjustment of security levels in biometric identification systems. Designed to address the increasing complexity due to diverse sensors and modalities, BioDSSL promotes system maintainability and resilience while ensuring a balance between usability and security for specific scenarios. Furthermore, it facilitates decentralization of biometric identification systems, by improving interoperability and abstraction. Decentralization inherently disperses the concentration of sensitive biometric data across various nodes, which could indirectly enhance privacy protection and limit the potential damage from localized security breaches. Therefore, BioDSSL is not just a technical improvement, but a step towards decentralized, resilient, and more secure biometric identification systems. This approach holds the promise of indirectly improving privacy while enhancing the reliability and adaptability of these systems amidst evolving threat landscapes and technological advancements.
Coordination of multi-agent systems has received significant attention during the past few years owing to its wide real-world applications, such as cooperative exploration, aircraft formation, and autonomous vehicle platooning. To address this issue, this research presents a novel method for multi-agent systems to navigate through environments with obstacles. The system consists of a group of agents with a leader-follower structure, where the leader aids in guiding the agents toward the target location and the followers are steered to maintain a flexible formation. To achieve cooperation, the agents communicate within a connected and undirected network, exchanging information within a specific radius. The leader's path is generated using the RRT* algorithm, which serves as a reference for the followers. A control law utilizes consensus and APF is then implemented, ensuring coordinated motion while maintaining safe distances among agents and between agents and obstacles. Finally, the effectiveness of the developed two-layer coordination strategy is verified by simulations.
Because data-driven AI-based methods can accommodate the intermittent nature of solar energy, they hold promise for forecasting solar Photovoltaic (PV) power generation. In order to determine which machine learning algorithm is the best effective in predicting the output of solar PV power, this study evaluates a number of well-constructed and optimized algorithms. In particular, the following four methods are examined: decision trees, random forests, extreme gradient boosting, and linear regression. These models are tested using the solar PV system that was put in place at the Applied Science Private University in Jordan. To compare the models equitably, three performance indicators are used: R-squared (R 2 ), Mean Absolute Error (MAE), and Root Mean Square Error (RMSE). The most successful model is the random forest one, with an RMSE of 11.04 kW and an MAE of 7.7 kW, and with a 98% R2 value. The robust prediction outcomes imply that improving the random forest model might increase prediction accuracy even more.
This article investigates the feasibility of implementing Decentralized Identity of Artifacts (DIDoA) on the BSN Spartan Blockchain. This research is part of a series of articles exploring the feasibility of DIDoA on various blockchain platforms. Securing the authenticity and ownership of artifacts is essential for protecting our cultural heritage. DIDoA leverages the inherent benefits of blockchain technology - trust, security and transparency - and offers a solution for creating tamper-proof records that track an artifact's history and ownership. To ensure consistent and comparable results across different blockchain networks, a standardized experimental setup, methodology, and processes are employed. This study utilizes a suite of tools and applications specifically designed to interact with the BSN Spartan Blockchain and its associated DIDoA functionalities (if available). The findings from this investigation will contribute to a broader understanding of DIDoA's suitability for various blockchain platforms and its potential impact on cultural heritage preservation efforts. We put the DIDoA use case to the test by implementing the building blocks of the framework - Decentralized Identifiers (DID) and Verifiable Credentials (VC) - on the blockchain. We conclude our study by analyzing the results and showing the advantages, potential challenges, limitations and future work.
The online solution of the optimization problems associated to the model predictive control (MPC) of constrained systems requires significant computational efforts. The Hindmarsh-Rose (HR) neuron model, which can simulate the functioning of biological neurons, is an example of a complex nonlinear system exhibiting chaotic behavior. Various control design methods have been applied to control the HR model, including nonlinear MPC based on online optimization. This paper proposes the design of an optimal controller for the HR model by using the NMPC approach based on deep neural networks (NMPC-DNN). The DNN of the controller is trained offline by using the optimal control trajectories obtained for a set of initial states of the HR model. Then online, the control input is computed simply by evaluating the function realized by the DNN. The advantages of the NMPC-DNN approach applied to the HR model are the significant reduction of the online computational complexity, as well as the simple software implementation of the controller. The performance of the NMPC-DNN controller designed for the HR model is studied by simulation experiments.
In this study, we utilize the mechanism of previous soft bio-inspired robots and propose a soft crawling robot that can move on the ground with various protrusions. The robot has two long flexible arms, which autonomously grasp protrusions and move by pulling and pushing the protrusions. The behavior of the robot is complex; nevertheless, it has only two motors. The complex behavior is realized by interacting with the environment. An actual robot was developed using silicone rubber, experiments were conducted and movements were realized.
Collecting the relevant list of patient phenotypes, known as deep phenotyping, can significantly improve the final diagnosis. As textual clinical reports are the richest source of phenotypes information, their automatic extraction is a critical task. The main challenges of this Information Extraction (IE) task are to identify precisely the text spans related to a phenotype and to link them unequivocally to referenced entities from a source such as the Human Phenotype Ontology (HPO). Recently, Language Models (LMs) have been the most suc-cessful approach for extracting phenotypes from clinical reports. Solutions such as PhenoBERT, relying on BERT or GPT, have shown promising results when applied to datasets built on the hypothesis that most phenotypes are explicitly mentioned in the text. However, this assumption is not always true in medical genetics. Hence, although the LMs carry powerful semantic abilities, their contributions are not clear compared to syntactic string-matching steps that are used within the current pipelines. The goal of this study is to improve phenotype extraction from clinical notes related to genetic diseases. Our contributions are threefold: First, we provide a clear definition of the phenotype extraction task from free text, along with a high-level overview of the involved functions. Second, we conduct an in-depth analysis of PhenoBERT, one of the best existing solutions, to evaluate the proportion of phenotypes predicted with simple string-matching. Third, we demonstrate how utilizing and incorporating large language models (LLMs) for span detection step can improve performance especially with implicit phenotypes. In addition, this experiment revealed that the annotations of existing dataset are not exhaustive, and that LLM can identify relevant spans missed by human labelers.