
The sudden increase in adoption of the Internet of Things (IoT) has revolutionized modern living but also brought unprecedented security challenges due to its distributed, heterogeneous, and resource-constrained nature. This review paper offers a comprehensive examination of machine learning (ML) and deep learning (DL) approaches tailored for intrusion detection and threat mitigation in IoT ecosystems. It explores the landscape of anomaly detection and classification techniques while analyzing their suitability, limitations, and deployment feasibility across IoT layers. The study also investigates the significance of feature engineering, model selection, and system scalability. A novel addition to this review is the integration of emerging trends such as explainable AI (XAI), which enhances transparency and trust in black-box ML/DL models, and federated learning (FL), a privacy-preserving paradigm that allows decentralized model training without raw data sharing. The synergy between FL and Edge AI is discussed to highlight real-time, low-latency security analytics at the network's edge. Comparative tables, domain-specific applications (e.g., smart homes, healthcare, and industrial IoT), and architectural illustrations support the discourse, providing readers with an up-to-date understanding of current capabilities and ongoing research challenges. This paper concludes with practical implications, research gaps, and future directions for building intelligent, secure, and explainable IoT security frameworks that respect user privacy and enable scalable deployment. This article is categorized under: Fundamental Concepts of Data and Knowledge > Explainable AI Technologies > Internet of Things Technologies > Machine Learning
Crowdsourcing has recently evolved as a distributed human problem-solving method and has received considerable interest from academics and practitioners in various domains. The proliferation of crowdsourcing has made it much simpler to utilize the intelligence and adaptability of many people to learn new knowledge to solve the problem of acquiring new knowledge. In the past, numerous crowdsourcing works have highlighted multiple aspects; however, no surveys have been conducted that focus on the entire crowdsourcing process. This concentrated survey provides a comprehensive review of the technical advances from a systematic perspective. This survey systematically reviews technical advances for a crowdsourcing process that contains four dimensions: task modeling, crowdsourcing data acquisition, the learning process, and predictive model learning, and proposes a comprehensive and scalable framework from CROWD4AI (Crowdsourcing Framework with 4 Dimensions for Artificial Intelligence). In addition, this paper focuses on each dimension's potential challenges and future direction, encouraging researchers to participate in crowdsourcing. To bridge theory with practice, we also include a detailed case study that demonstrates the real-world application of our proposed framework in the context of annotating cultural heritage damages using crowdsourced input. The case study illustrates how the framework supports effective task design, label collection, robust learning strategies, and accurate predictive modeling in a practical setting.This article is categorized under: Technologies > Crowdsourcing Technologies > Machine Learning
Managing complex disaster risks requires interdisciplinary efforts. Breaking down silos between law, social sciences, and natural sciences is critical for all processes of disaster risk reduction. It is essential to explore how AI enhances understanding of legal frameworks and environmental management, while also examining how legal and environmental factors may limit AI's role in society. From a co-production review perspective, drawing on insights from lawyers, social scientists, and environmental scientists, principles for responsible data mining are proposed based on safety, transparency, fairness, accountability, and contestability. This discussion offers a blueprint for interdisciplinary collaboration to create adaptive law systems based on AI integration of knowledge from environmental and social sciences. When social networks are useful for mitigating disaster risks based on AI, the legal implications related to privacy and liability of the outcomes of disaster management must be considered. Fair and accountable principles emphasize environmental considerations and foster socioeconomic discussions related to public engagement. AI also has an important role to play in education, bringing together the next generations of law, social sciences, and natural sciences to work on interdisciplinary solutions in harmony. Although emerging AI approaches can be powerful tools for disaster management, they must be implemented with ethical considerations and safeguards to address concerns about bias, transparency, and privacy. The responsible execution of AI approaches, based on the dynamic interplay between AI, law, and environmental risk, promotes sustainable and equitable practices in data mining.
Early diagnosis of abnormal cervical cells enhances the chance of prompt treatment for cervical cancer (CrC). Artificial intelligence (AI)-assisted decision support systems for detecting abnormal cervical cells are developed because manual identification needs trained healthcare professionals, and can be difficult, time-consuming, and error-prone. The purpose of this study is to present a comprehensive review of AI technologies used for detecting cervical pre-cancerous lesions and cancer. The review study includes studies where AI was applied to Pap Smear test (cytological test), colposcopy, sociodemographic data and other risk factors, histopathological analyses, magnetic resonance imaging-, computed tomography-, and positron emission tomography-scan-based imaging modalities. We performed searches on Web of Science, Medline, Scopus, and Inspec. The preferred reporting items for systematic reviews and meta-analysis guidelines were used to search, screen, and analyze the articles. The primary search resulted in identifying 9745 articles. We followed strict inclusion and exclusion criteria, which include search windows of the last decade, journal articles, and machine/deep learning-based methods. A total of 58 studies have been included in the review for further analysis after identification, screening, and eligibility evaluation. Our review analysis shows that deep learning models are preferred for imaging techniques, whereas machine learning-based models are preferred for sociodemographic data. The analysis shows that convolutional neural network-based features yielded representative characteristics for detecting pre-cancerous lesions and CrC. The review analysis also highlights the need for generating new and easily accessible diverse datasets to develop versatile models for CrC detection. Our review study shows the need for model explainability and uncertainty quantification to increase the trust of clinicians and stakeholders in the decision-making of automated CrC detection models. Our review suggests that data privacy concerns and adaptability are crucial for deployment hence, federated learning and meta-learning should also be explored. This article is categorized under: Fundamental Concepts of Data and Knowledge > Explainable AI Technologies > Machine Learning Technologies > Classification
Pests pose a major danger to a variety of industries, including agriculture, public health, and ecosystems. Fast and precise pest detection, as well as the ability to predict infestations, are required for effective pest management tactics. This paper provides a comprehensive literature review on this subject to provide an overview of the state of research on pest detection and infestation prediction. The paper investigates and presents background information on the necessity of pest control as well as the difficulty in recognizing pests and forecasting. Several strategies, including approaches to data collection, modeling, and assessment of models, are reviewed in the research described. The authors examine various pest detection methods involving the utilization of convolutional neural networks and several object detection architectures categorized broadly into one-stage and two-stage object detection algorithms. Methods for predicting pest infestations that involve regression, classification, and time series forecasting are also thoroughly investigated. The challenges of recognizing pests and predicting infestations are underlined, as are issues with data quality, feature selection, and model interpretability. The report also indicates the limitations to pest detection and infestation prediction as well as intriguing topics for further research on the same. The findings of the literature research demonstrate how Artificial Intelligence, Computer Vision, and the Internet of Things have been applied for Pest Detection and Infestation Prediction. The research serves as a base for surveying and summarizing the approaches utilized for the task of pest detection (an object detection problem) and pest infestation prediction (a forecasting problem) and its findings and recommendations serve as a platform for future study and the development of effective pest management solutions.
The application of machine learning techniques in the field of tourism is experiencing a remarkable growth, as they allow to propose efficient solutions to problems present in this sector, by means of an intelligent analysis of data in their specific context. The increase of work in this field requires an exhaustive analysis through a quantitative approach of research activity, contributing to a deeper understanding of the progress of this field. Thus, different approaches in the field of tourism will be analyzed, such as planning, forecasting, recommendation, prevention, and security, among others. As a result of this analysis, among other findings, the greater impact of supervised learning in the field of tourism, and more specifically those techniques based on neural networks, has been confirmed. The results of this study would allow researchers not only to have the most up-to-date and accurate overview of the application of machine learning in tourism, but also to identify the most appropriate techniques to apply to their domain of interest, as well as other similar approaches with which to compare their own solutions. This article is categorized under: Application Areas > Society and Culture Technologies > Machine Learning Application Areas > Business and Industry
Explainability in the field of event detection is a new emerging research area. For practitioners and users alike, explainability is essential to ensuring that models are widely adopted and trusted. Several research efforts have focused on the efficacy and efficiency of event detection. However, a human-centric explanation approach to existing event detection solutions is still lacking. This paper presents an overview of a conceptual framework for human-centric semantic-based explainable event detection with the acronym HUSEED. The framework considered the affordances of XAI and semantics technologies for human-comprehensible explanations of events to facilitate 5W1H explanations (Who did what, when, where, why, and how). Providing this kind of explanation will lead to trustworthy, unambiguous, and transparent event detection models with a higher possibility of uptake by users in various domains of application. We illustrated the applicability of the proposed framework by using two use cases involving first story detection and fake news detection. A conceptual framework for human-centric and semantic-based explainable event detection (HUSEED).image
Design and development of new drug molecules are essential for the survival of human society. New drugs are designed for therapeutic purposes to combat new diseases. Besides treating new diseases, new drug development is also needed to treat pre-existing diseases more effectively and reduce the existing drugs' side effects. The design of drugs involves several steps, from the discovery of the drug molecule to its commercialization in the market. One of the most critical steps in drug design is to find the molecular interactions between the target (infected) molecule and the drug molecule. Several complex chemical equations need to be solved to determine the molecular interactions. In the late 20th Century, the advancement of computational technologies has made the solution of chemical equations relatively easier and faster. Moreover, the design of drug molecules involves multi-criteria optimization. Classical computational methodologies have been used for drug design since the end of the 20th Century. However, nowadays, more advanced computational methodologies are inevitable in designing drugs for new diseases and drugs with fewer side effects. In this context, the quantum computing paradigm has proved beneficial in drug design due to its advanced computational capabilities. This paper presents a state-of-the-art comprehensive review of the quantum computing-based methodologies involved in drug design. A comparative study is made about the different quantum-aided drug design methods, stating each methodology's merits and demerits. The review work presented in this manuscript will help new researchers assess the present state-of-the-art concept of quantum-based drug design. This article is categorized under: Technologies > Structure Discovery and Clustering Technologies > Computational Intelligence Application Areas > Health Care
Due to advancements in data collection, storage, and processing techniques, machine learning has become a thriving and dominant paradigm. However, one of its main shortcomings is that the classical machine learning paradigm acts in isolation without utilizing the knowledge gained through learning from related tasks in the past. To circumvent this, the concept of Lifelong Machine Learning (LML) has been proposed, with the goal of mimicking how humans learn and acquire cognition. Human learning research has revealed that the brain connects previously learned information while learning new information from a single or small number of examples. Similarly, an LML system continually learns by storing and applying acquired information. Starting with an analysis of how the human brain learns, this paper shows that the LML framework shares a functional structure with the brain when it comes to solving new problems using previously learned information. It also provides a description of the LML framework, emphasizing its similarities to human brain learning. It also provides citation graph generation and scientometric analysis algorithms for the LML literatures, including information about the datasets and evaluation metrics that have been used in the empirical evaluation of LML systems. Finally, it presents outstanding issues and possible future research directions in the field of LML.This article is categorized under: Technologies > Machine Learning
From the latter half of the last decade, there has been a growing interest in developing algorithms for automatically solving mathematical word problems (MWP). It is a challenging and unique task that demands blending surface level text pattern recognition with mathematical reasoning. In spite of extensive research, we still have a lot to explore for building robust representations of elementary math word problems and effective solutions for the general task. In this paper, we critically examine the various models that have been developed for solving word problems, their pros and cons and the challenges ahead. In the last 2 years, a lot of deep learning models have recorded competing results on benchmark datasets, making a critical and conceptual analysis of literature highly useful at this juncture. We take a step back and analyze why, in spite of this abundance in scholarly interest, the predominantly used experiment and dataset designs continue to be a stumbling block. From the vantage point of having analyzed the literature closely, we also endeavor to provide a road-map for future math word problem research. This article is categorized under: Technologies > Machine Learning Technologies > Artificial Intelligence Fundamental Concepts of Data and Knowledge > Knowledge Representation
The growing popularity of wearable health devices like fitness trackers and smartwatches enables continuous personal health monitoring but also raises significant privacy concerns due to the real-time collection of sensitive data. Many users are unaware of vulnerabilities that could lead to unauthorized access or discrimination if health information is revealed without consent. However, even informed users may willingly share data despite understanding privacy risks. The recent implementation of the General Data Protection Regulation (GDPR) in the EU and states taking initiatives to regulate privacy shows growing regulatory efforts to address these threats. This paper evaluates the key privacy threats posed specifically by consumer wearable devices. It provides a focused analysis of how health data could be exploited or shared without users' knowledge and the security flaws that enable such risks. Potential solutions including improving protections, empowering user control, enhancing transparency, and strengthening regulations are examined. However, it is argued that effective change requires balancing privacy risks with health benefits while also considering human decision-making behaviors. The paper concludes by proposing a multifaceted approach to enable informed choices about wearable health data. This article is categorized under: Application Areas > Health Care Commercial, Legal, and Ethical Issues > Fairness in Data Mining Commercial, Legal, and Ethical Issues > Legal Issues
The timely identification of significant memory concern (SMC) is crucial for proactive cognitive health management, especially in an aging population. Detecting SMC early enables timely intervention and personalized care, potentially slowing cognitive disorder progression. This study presents a state-of-the-art review followed by a comprehensive evaluation of machine learning models within the randomized neural networks (RNNs) and hyperplane-based classifiers (HbCs) family to investigate SMC diagnosis thoroughly. Utilizing the Alzheimer's Disease Neuroimaging Initiative 2 (ADNI2) dataset, 111 individuals with SMC and 111 healthy older adults are analyzed based on T1W magnetic resonance imaging (MRI) scans, extracting rich features. This analysis is based on baseline structural MRI (sMRI) scans, extracting rich features from gray matter (GM), white matter (WM), Jacobian determinant (JD), and cortical thickness (CT) measurements. In RNNs, deep random vector functional link (dRVFL) and ensemble dRVFL (edRVFL) emerge as the best classifiers in terms of performance metrics in the identification of SMC. In HbCs, Kernelized pinball general twin support vector machine (Pin-GTSVM-K) excels in CT and WM features, whereas Linear Pin-GTSVM (Pin-GTSVM-L) and Linear intuitionistic fuzzy TSVM (IFTSVM-L) performs well in the JD and GM features sets, respectively. This comprehensive evaluation emphasizes the critical role of feature selection, feature based-interpretability and model choice in attaining an effective classifier for SMC diagnosis. The inclusion of statistical analyses further reinforces the credibility of the results, affirming the rigor of this analysis. The performance measures exhibit the suitability of this framework in aiding researchers with the automated and accurate assessment of SMC. The source codes of the algorithms and datasets used in this study are available at . This article is categorized under: Technologies > Classification Technologies > Machine Learning Application Areas > Health Care
With proliferation of Big Data, organizational decision making has also become more complex. Business Intelligence (BI) is no longer restricted to querying about marketing and sales data only. It is more about linking data from disparate applications and also churning through large volumes of unstructured data like emails, call logs, social media, News, and so on in an attempt to derive insights that can also provide actionable intelligence and better inputs for future strategy making. Semantic technologies like knowledge graphs have proved to be useful tools that help in linking disparate data sources intelligently and also enable reasoning through complex networks that are created as a result of this linking. Over the last decade the process of creation, storage, and maintenance of knowledge graphs have sufficiently matured, and they are now making inroads into business decision making also. Very recently, these graphs are also seen as a potential way to reduce hallucinations of large language models, by including these during pre-training as well as generation of output. There are a number of challenges also. These include building and maintaining the graphs, reasoning with missing links, and so on. While these remain as open research problems, we present in this article a survey of how knowledge graphs are currently used for deriving business intelligence with use-cases from various domains. This article is categorized under: Algorithmic Development > Text Mining Application Areas > Business and Industry
Smartphones and personal sensing technologies have made collecting data continuously and in real time feasible. The promise of pervasive sensing technologies in the realm of mental health has recently garnered increased attention. Using Artificial Intelligence methods, it is possible to forecast a person's emotional state based on contextual information such as their current location, movement patterns, and so on. As a result, conditions like anxiety, stress, depression, and others might be tracked automatically and in real-time. The objective of this research was to survey the state-of-the-art autonomous psychological health monitoring (APHM) approaches, including those that make use of sensor data, virtual chatbot communication, and artificial intelligence methods like Machine learning and deep learning algorithms. We discussed the main processing phases of APHM from the sensing layer to the application layer and an observation taxonomy deals with various observation devices, observation duration, and phenomena related to APHM. Our goal in this study includes research works pertaining to working of APHM to predict the various mental disorders and difficulties encountered by researchers working in this sector and potential application for future clinical use highlighted.
Future connected and autonomous vehicles (CAVs) must be secured against cyberattacks for their everyday functions on the road so that safety of passengers and vehicles can be ensured. This article presents a holistic review of cybersecurity attacks on sensors and threats regarding multi‐modal sensor fusion. A comprehensive review of cyberattacks on intra‐vehicle and inter‐vehicle communications is presented afterward. Besides the analysis of conventional cybersecurity threats and countermeasures for CAV systems, a detailed review of modern machine learning, federated learning, and blockchain approach is also conducted to safeguard CAVs. Machine learning and data mining‐aided intrusion detection systems and other countermeasures dealing with these challenges are elaborated at the end of the related section. In the last section, research challenges and future directions are identified.This article is categorized under:Commercial, Legal, and Ethical Issues > Security and PrivacyTechnologies > Machine LearningTechnologies > Internet of Things
Several open-source programming languages, particularly R and Python, are utilized in industry and academia for statistical data analysis, data mining, and machine learning. While most commercial software programs and programming languages provide a single way to deliver a statistical procedure, open-source programming languages have multiple libraries and packages offering many ways to complete the same analysis, often with varying results. Applying the same statistical method across these different libraries and packages can lead to entirely different solutions due to the differences in their implementations. Therefore, reliability and accuracy should be essential considerations when making library and package usage decisions while conducting statistical analysis using open source programming languages. Instead, most users take this for granted, assuming that their chosen libraries and packages produce accurate results for their statistical analysis. To this extent, this study assesses the estimation accuracy and reliability of Python and R's various libraries and packages by evaluating the univariate summary statistics, analysis of variance (ANOVA), and linear regression procedures using benchmarking data from the National Institutes of Standards and Technology (NIST). Further, experimental results are presented comparing machine learning methods for classification and regression. The libraries and packages assessed in this study include the stats package in R and Pandas, Statistics, NumPy, statsmodels, SciPy, statsmodels, scikit-learn, and pingouin in Python. The results show that the stats package in R and statsmodels library in Python are reliable for univariate summary statistics. In contrast, Python's scikit-learn library produces the most accurate results and is recommended for ANOVA. Among the libraries and packages assessed for linear regression, the results demonstrated that the stats package in R is more reliable, accurate, and flexible; thus, it is recommended for linear regression analysis. Further, we present results and recommendations for machine learning using R and Python. This article is categorized under: Algorithmic Development > Statistics Application Areas > Data Mining Software Tools
Ultra-precision machining (UPM), one of the most advanced machining techniques that can produce exact components, significantly impacts the technological community. The significance of UPM attracts the attention of academic and industrial partners. As a result of the rapid development of UPM caused by technological advancement, it is necessary to revisit the current stages and evolution of UPM to sustain and advance this technology. The state of the art in UPM is first investigated systematically in this study by identifying the current four major UPM themes. The UPM thematic network is then built, along with a structural analysis of the network, to determine the interactions between each theme and the primary roles of theme members responsible for the interactions. Furthermore, the "bridge" role is assigned to the specific UPM theme content. On the other hand, Sentiment analysis is conducted to determine how the academic community at UPM feels about the themes for UPM research to focus on those themes with a need for more confidence. Considering the above findings, the future perspective of UPM and suggestions for its advancement are discussed and provided. This study provides a comprehensive understanding and the current state-of-the-art review of UPM technology by a text mining technique to critically analyze its research content, as well as suggestions to enhance UPM development by focusing on its current challenges, thereby assisting academia and institutions in leveraging this technology to benefit society.
Satellite image time-series are time series produced from remote sensing images; they generally correspond to features or indicators extracted from those images. With the increasing availability of remote sensing images and new methodologies to process such data, image time-series methods have been used extensively for assessing temporal pattern detection, monitoring, classification, object detection, and feature estimation. Since the study of time series is broad, this article focuses on analyzing articles related to forecasting the value of one or more attributes of the image time-series. The image time series forecasting (ITSF) problem appears in different disciplines; most focus on improving the quality of life by harnessing natural resources for sustainable development and minimizing the lethality of dangerous natural phenomena. Scientists tackle these problems using different tools or methods depending on the application. This review analyzes the field's leading, most recent contributions, grouping them by application area and solution methods. Our findings indicate that artificial neural networks, regression trees, support vector regression, and cellular automata are the most common methods for ITSF. Application areas address this problem as renewable energy, agriculture, and land-use change. This study retrieved and analyzed relevant information about the recent activity of image time series forecasting, generating a reproducible list of the most pertinent articles in the field published from 2009 to 2021. To the author's best knowledge, this is the first review presenting and analyzing a reproducible list of the most relevant state-of-the-art articles focusing on the applications, techniques, and research trends for ITSF. This article is categorized under: Algorithmic Development > Spatial and Temporal Data Mining Technologies > Machine Learning Technologies > Prediction
This article gives a brief overview of various aspects of data mining of multispectral image data. We focus on specifically the remote sensing satellite images acquired using multispectral imaging (MSI), given the technology used across multiple knowledge domains, such as chemistry, medical imaging, remote sensing, and so on with a sufficient amount of variation. In this article, the different data mining processes are reviewed along with state‐of‐the‐art methods and applications. To study data mining, it is important to know how the data are acquired and preprocessed. Hence, those topics are briefly covered in the article. The article concludes with applications demonstrating the knowledge discovery from data mining, modern challenges, and promising future directions for MSI data mining research. This article is categorized under: Application Areas > Science and Technology Fundamental Concepts of Data and Knowledge > Knowledge Representation Fundamental Concepts of Data and Knowledge > Big Data Mining
In order to engineer new materials, structures, systems, and processes that address persistent challenges, engineers seek to tie causes to effects and understand the effects of causes. Such a pursuit requires a causal investigation to uncover the underlying structure of the data generating process (DGP) governing phenomena. A causal approach derives causal models that engineers can adopt to infer the effects of interventions (and explore possible counterfactuals). Yet, and for the most part, we continue to design experiments in the hope of empirically observing engineered intervention(s). Such experiments are idealized, complex, and costly and hence are narrow in scope. On the contrary, a causal investigation will allow us to peek into the how and why of a DGP and provide us with the essential means to articulate a causal model that accurately describes the phenomenon on hand and better predicts the outcome of possible interventions. Adopting a causal approach in engineering is perhaps more warranted than ever-especially with the rise of big data and the adoption of artificial intelligence (AI); wherein AI models are naivety presumed to describe causal ties. To bridge such knowledge gap, this primer presents fundamental principles behind causal discovery, causal inference, and counterfactuals from an engineering perspective.