The increased deployment of Internet of Things (IoT) devices has led to greater complexity and scale in network traffic, introducing challenges in effectively detecting cyber threats. This study compares two network flow feature extraction tools, NFStream and Tranalyzer, through their performance in detecting malicious network traffic using a Multilayer Perceptron (MLP) neural network. The evaluation was conducted using datasets generated from laboratory-based IoT setups involving edge devices, such as the Raspberry Pi 4 and Jetson Nano, as well as data from the publicly available ToN-IoT dataset. Each tool independently processes the same raw network captures, creating two distinct feature sets. Both datasets underwent identical preprocessing steps, including balanced undersampling, categorical encoding, and feature scaling, to ensure fair comparison. MLP neural networks, implemented using the Keras Functional API, were trained, validated, and tested on each dataset separately. The results, represented by global and network-specific confusion matrices, demonstrate the high accuracy of both tools, with subtle but measurable differences in performance. Both NFStream and Tranalyzer demonstrated strong and consistent performance under identical evaluation settings, confirming their capability to support machine learning-based detection of cyberattacks in IoT environments.
The rapid expansion of Internet of Things (IoT) devices has significantly increased the attack surface, necessitating the development of robust and privacy-preserving intrusion detection systems. This research delves into the application of Federated Learning (FL) to detect network-based attacks in IoT environments, utilizing the NF-ToN-IoT dataset. We compare a centralized machine learning model with a federated counterpart, employing the Federated Averaging (FedAvg) algorithm. Both models employ a Multi-Layer Perceptron (MLP) architecture trained on NetFlow features for binary classification of benign versus malicious traffic. Our experimental results reveal that the federated model achieves comparable and even slightly superior performance across various metrics, including precision, recall, and F1-score, while preserving data privacy by decentralizing the training data. These findings underscore the potential of FL as a viable alternative to traditional intrusion detection systems in real-world, privacy-sensitive IoT scenarios.
The rapid growth of Internet of Things (IoT) deployments has intensified cybersecurity challenges, as machine-learning-based intrusion detection often generalizes poorly across heterogeneous network environments. This study evaluates cross-domain generalization strategies for IoT cyberattack detection using NFStream-based flow features. Six learning configurations were assessed across two heterogeneous datasets, ToN-IoT and UBU-LAB, and three supervised architectures: a Multilayer Perceptron (MLP), a one-dimensional Convolutional Neural Network (CNN-1D), and an FT-Transformer. The configurations comprise single-domain training in each domain, transfer learning in both directions, centralized multi-domain training, and federated multi-domain training, with results reported as mean and standard deviation over five folds. All architectures achieved strong in-domain performance, with accuracy and F1-score above 97.5% on ToN-IoT and at least 99.95% on UBU-LAB. However, direct cross-domain evaluation caused severe degradation across all architectures, confirming that architecture replacement alone does not prevent domain-shift effects. Transfer learning recovered target-domain performance, although source-domain retention proved architecture-dependent. Centralized training yielded the most stable multi-domain behavior, whereas federated learning offered a privacy-preserving alternative whose performance depended more strongly on the architecture, with CNN-1D achieving the most balanced results. These findings indicate that explicit generalization strategies are more relevant than architecture replacement alone when selecting IoT cyberattack detection approaches under accuracy, privacy, and data-sharing constraints.
Automatic analysis of neonatal facial behavior in real Neonatal Intensive Care Unit (NICU) environments is challenged by high data variability, visual noise, and the need for interpretable models suitable for clinical use. Facial expressions are a core component of neonatal pain, discomfort, and alertness assessments; however, their evaluation remains largely subjective in routine practice. In this study, a dataset of neonatal video recordings acquired under routine NICU conditions was analyzed, comprising 991 annotated videos after preprocessing. An attention-based convolutional neural network was employed for the frame-level binary classification of three clinically relevant facial traits: eye opening, eye frowning, and mouth opening. To address interpretability requirements in safety-critical settings, a dual explainability framework combining intrinsic attention mechanisms and post-hoc SHAP analysis was adopted. The proposed approach achieved strong performance, particularly for eye-related traits, despite challenging acquisition conditions. Explainability analyses showed that predictions were driven by anatomically meaningful facial regions, supporting model transparency and trustworthiness for clinical deployment.
This data article presents a labelled flow-based network traffic dataset collected from a controlled Internet of Things (IoT) laboratory environment. The dataset captures network communication generated by Raspberry Pi-based IoT nodes configured to emulate service and client roles. Traffic was recorded during normal operations and during the execution of predefined cyberattack scenarios within an isolated experimental network.Network traffic was recorded at the packet level using passive network monitoring and stored in PCAPNG format. The packet captures were subsequently processed into bidirectional network flows, producing flow records with statistical and temporal attributes derived from the observed packet exchanges.Cyberattack-related flows were labelled using the experimental ground-truth markers recorded during each attack campaign, complemented by the fixed attacker node IP address. Flows outside the marked intervals were labelled as benign and corresponded to regular device communication. This combined labelling approach reduces the potential for overlap between benign and attack activities.The dataset covers nine attack scenarios grouped into six attack categories. It is released through a structured repository containing raw packet captures, labelled flow files, and supporting metadata for flow-based IoT traffic analysis, cyberattack detection research, and optional re-labelling.
Quality assurance stands as a pivotal phase across all manufacturing processes, particularly within the textile industry. Presently, textile inspections heavily rely on human visual assessment due to deficiencies in commercial solutions. In response, this research advocates for the adoption of Deep Learning models to automate fabric quality control within dynamic and contemporary production environments. Specifically, Convolutional Neural Networks are scrutinized using authentic images sourced from the Batavia and Sarga weave production lines. The experimental results identify DenseNet121 and Incep-tionV3 as the most effective models for Batavia and Sarga weaves, respectively. DenseNet121 demonstrates balanced performance across key metrics for Batavia weave, while InceptionV3 excels in Sarga weave, particularly in F1-score and AU-ROC. These findings underscore the potential of DL models to enhance the accuracy and efficiency of textile quality control.
Flow-based network traffic analysis is widely adopted for securing IoT environments owing to its scalability and computational efficiency. However, reliance on basic statistical flow features may fail to detect subtle layer 2 intrusions and can lead to laboratory overfitting when static identifiers are implicitly learned. This paper presents a rigorous evaluation on a publicly available IoT flow dataset comprising seven attack categories, using a methodology explicitly designed to prevent data leakage. While tree-based ensemble models achieved near-perfect detection (F1 > 0.99) for most attacks, initial experiments revealed a critical performance gap for ARP Spoofing. Even after hyperparameter optimization, standard flow statistics yielded a Recall of 0.86, leaving a significant portion of malicious activity undetected. Through Explainable AI (SHAP), we demonstrated the statistical overlap that causes this limitation. Finally, we propose a context-aware feature engineering approach—specifically introducing flow symmetry ratios and cardinality metrics—which successfully restored detection capabilities to a perfect F1-Score of 1.0, proving the viability of behavioral state analysis for complex IoT threats.
In this study, the impact of various data augmentation techniques on defect detection in Batavia textiles was analyzed using a publicly available dataset consisting of high-resolution images from textile manufacturing processes. A baseline model trained without augmentation was compared with models trained using both individual and combined augmentation techniques, including horizontal and vertical flips, rotations, ZCA whitening, and brightness adjustments. Experiments were conducted using 5-fold stratified cross-validation, and the performance of the convolutional neural network DenseNet121 was evaluated using metrics suited for imbalanced binary classification tasks. The results show that not all augmentation strategies led to substantial improvements; while certain combinations of augmentation techniques resulted in the best performance-time balance, other transformations either had minimal impact or even degraded performance. Furthermore, the obtained results proof that AU-ROC remains stable across most simple transformations, suggesting that data augmentation, while useful, does not always yield a noticeable improvement in model performance.
In industrial environments, the early detection of anomalies in equipment, such as hydrogen compressors, is critical for ensuring operational safety and reliability. This paper proposes a novel hybrid approach that combines time series forecasting with supervised classification to predict future alarm and warning events. The study uses real-world data from hydrogen compressors provided by Hiperbaric, a company specialized in industrial equipment for high-pressure technologies. The data underwent extensive preprocessing to ensure quality and temporal consistency. Eight tailored datasets were constructed to represent different components and types of alerts. The methodology integrates a Vector Autoregressive model for forecasting endogenous variables and an XGBoost classifier for predicting binary alarm signals. To address the inherent class imbalance in the data, a comprehensive evaluation was conducted using metrics such as AU-ROC, AU-PR, F1-score, and G-mean. The results show that the proposed approach effectively captures temporal dependencies and enables the accurate early detection of alarms and warnings. This study demonstrates the practical applicability of combining forecasting and classification models for the predictive maintenance of hydrogen compression systems.
This paper presents a novel iterative algorithm for determining contact-rate parameters in individual-based propagation models. Conventional mathematical epidemiological approaches frequently utilize compartmental models with uniform epidemiological coefficients, which are unable to account for the heterogeneity of individual behaviors and interactions. This research employs Physics Informed Neural Networks (PINNs), which have demonstrated efficacy in addressing differential equations and parameter estimation in both direct and inverse problems. PINNs are applied to individual-based models, which consider unique epidemiological characteristics for each individual or device, thus addressing the limitations of global models. The proposed method is particularly pertinent to the modeling of malware propagation in Internet of Things (IoT) networks, where each device exhibits distinct interaction patterns. The study employs a cellular automaton framework to simulate malware spread, thereby demonstrating the potential of PINNs in accurately estimating time-dependent contact rates. The methodology incorporates specific infection probabilities and recovery rates, thereby demonstrating its applicability to complex, heterogeneous systems. This approach represents a significant advancement in our understanding and prediction of the dynamics of disease and malware propagation, offering a robust tool for both biological and digital epidemiology. The findings indicate that individual-based models, enhanced with PINN-driven parameter estimation, can more accurately reflect real-world scenarios, thereby facilitating the development of more effective strategies for controlling epidemics and mitigating cybersecurity threats.
Websites, as essential digital assets, are highly vulnerable to cyberattacks because of their high traffic volume and the significant impact of breaches. This study aims to enhance the identification of web traffic attacks by leveraging machine learning techniques. A methodology was proposed to extract relevant features from HTTP traces using the CSIC2010 v2 dataset, which simulates e-commerce web traffic. Ensemble methods, such as Random Forest (RF) and Extreme Gradient Boosting, were employed and compared against baseline classifiers, including k-nearest Neighbor, LASSO, and Support Vector Machines. The results demonstrate that the ensemble methods outperform baseline classifiers by approximately 20% in predictive accuracy, achieving an Area Under the ROC Curve of 0.989. Feature selection methods such as Information Gain, LASSO, and RF further enhance the robustness of these models. This study highlights the efficacy of ensemble models in improving attack detection while minimizing performance variability, offering a practical framework for securing web traffic in diverse application contexts.
This article presents a multivariate time series dataset detailing the physicochemical degradation of an industrial metalworking fluid (MWF). The data were collected continuously over several months from a test tank under typical operational conditions at an industrial facility in Spain. Four critical variables were monitored using industrial-grade sensors: pH, temperature, concentration, and conductivity. The dataset is provided in five CSV files. The primary file, measures.csv, contains the preprocessed time series at a uniform 5-minute frequency, with authentic missing data gaps intentionally preserved to reflect real-world sensor and connectivity issues. The four additional files serve as a comprehensive benchmark for data imputation algorithms. Each of these benchmark files corresponds to a single variable and includes the original data alongside imputed values generated by five distinct methods: K-Nearest Neighbours (KNN), a hybrid model (HybridKCL), an LSTM-based Variational Autoencoder (LSTM-VAE), and both pre-trained and fine-tuned versions of the MOMENT foundation model. This resource enables researchers and practitioners to develop, validate, and compare predictive maintenance models, anomaly detection systems, and advanced imputation techniques. Furthermore, it serves as a valuable educational tool for addressing common challenges in industrial IoT data, fostering advancements in sustainable and efficient manufacturing.
This dataset, collected during November 2022 at Textil Santanderina, a leading textile manufacturer based in Cabezón de la Sal (Cantabria, Spain), comprises high-resolution images of Batavia and Sarga fabrics. The images were captured as part of a project to document and analyze the intricate weaves and patterns of these fabrics. Using a high-resolution camera under controlled lighting conditions, detailed images were obtained to ensure consistent quality and accurate representation of the fabric's texture and colour. The dataset is provided in processed format, where images have been downscaled from 16 bits to 8 bits, cropped, and classified into cases and controls. The primary reuse potential of this dataset lies in its application for Artificial Intelligence (AI) and Machine Learning (ML) models aimed at defect detection in textile manufacturing. By leveraging these high-quality processed images, researchers and developers can train models to identify and classify various types of fabric defects, such as weave inconsistencies, colour variations, and surface irregularities. This can significantly enhance the efficiency and accuracy of quality control processes in textile production. Additionally, the dataset serves as a valuable resource for academic research in textile engineering and material science. It can be used to study the properties and behaviours of Batavia and Sarga weaves under different conditions, contributing to advancements in fabric design and manufacturing techniques. The detailed visual information provided by the processed images also supports the development of new methodologies for automated textile inspection and quality assurance. By making this dataset available, Textil Santanderina and University of Burgos aim to support innovation and improvement in textile quality control through AI-driven solutions, fostering collaboration and development within the industry.
Metalworking fluids (MWFs) are essential in machining processes, although their degradation over time requires continuous monitoring to prevent operational and economic losses. This study addresses missing data in a MWF sensor time series, which hinders the reliability of predictive maintenance. Advanced imputation methods (including pre-trained and fine-tuned MOMENT-1, LSTM-VAE, KNN, and HybridKCL) were evaluated for reconstructing gaps in four critical MWF properties: pH, temperature, concentration, and conductivity. The performance was quantitatively assessed using the MAE and RMSE on artificial masked data. Among the evaluated methods, the fine-tuned MOMENT-1 model generally outperformed the other methods across the variables, exhibiting a favorable balance of low reconstruction error and high visual consistency. However, qualitative inspection remains essential to verify the plausibility of the reconstructed dynamics. These findings contribute to improving the integrity of MWF monitoring data, enabling more reliable predictive analytics, and supporting efficient and sustainable manufacturing.
center dot Automated defect detection in fabrics is a key challenge in quality control within the textile industry. This study proposes a deep learning-based methodology to identify defects in Batavia and Sarga fabrics. In the first stage, an autoencoder was used to filter anomalous images, enabling the creation of a dataset with sufficient defective cases, which are otherwise difficult to obtain in textile production. Subsequently, convolutional neural networks (DenseNet121, EfficientNetB0/B3, Xception, and VGG) were trained using data augmentation techniques and stratified cross-validation. For Batavia fabrics, DenseNet121 achieved an AU-ROC of 0.88 and an AU-PR of 0.93, demonstrating high detection capability. For Sarga fabrics, three different references (42402, 45433, and 43105) were considered, showing more variable performance across models and datasets. Nonetheless, models such as ResNet101 and Xception achieved competitive results. The results indicate that the combination of autoencoder and CNN facilitates the generation of balanced datasets and enables consistent defect detection, although performance depends on the type of fabric and the specific reference, suggesting that model selection should be adapted to the characteristics of each case.
The safe and efficient operation of hydrogen refueling stations is essential to support the global transition towards low-carbon energy systems. However, the scarcity of real-world operational data remains a major obstacle for advanced monitoring and anomaly detection. This study proposes a deep learning framework that combines LSTM-based synthetic data generation with unsupervised anomaly detection. A generative LSTM was used to simulate 756 realistic hydrogen refueling scenarios enriched with physically plausible anomalies. An LSTM-Autoencoder was subsequently trained to detect deviations in key process variables, achieving 92
Predicting neurodevelopmental outcomes in very preterm infants is critical, but clinically-acquired data like longitudinal total brain volume (TBV), as macroscopic index of brain growth, are often sparse and irregular, hindering accurate prognosis. We studied 294 very preterm infants with TBV measured longitudinally and neurodevelopmental outcomes assessed at 2 years (Bayley-III) and 8 years (WISC-V). To handle missingness and irregular sampling, we compared six imputation strategies (mean, MissForest, MICE, GP, MGP, and MGP (initGP)) and trained a semi-supervised classifier on the imputed TBV trajectories. Performance was evaluated using multiple classification metrics, and results were summarized by averaging across outcomes and ages of study. Across analyses, a novel variant of Missing Gaussian Process (MGP) initialized with a GP fit (MGP (initGP)), which leverages individual patient trajectories to stabilize estimates, achieved the best average performance. Its advantages were most consistent and significant for 8-year outcomes, highlighting its strength in modeling longer developmental trajectories. While performance at 2 years was more modest, this likely reflects the intrinsic challenges of early-term prediction from TBV alone. We therefore recommend MGP (initGP) as a strong default for imputing longitudinal TBV for prognostic studies.
Continuous monitoring of neonatal behavior in the Neonatal Intensive Care Unit (NICU) is essential for early detection of neurological disorders. Among behavioral indicators, eye state (open vs. closed) serves as a clinically relevant marker for alertness, sedation, and responsiveness. This study presents a deep learning-based system for automated eye state detection in NICU video recordings. Using a manually labeled dataset of 7,388 facial frames extracted from 154 clinical videos, we trained and evaluated binary classifiers based on VGG16 and VGG19 convolutional neural network architectures. A five-fold cross-validation scheme was implemented to ensure subject-independent evaluation. The models achieved mean frame-level accuracies above 0.86 and AUC-PR values of 0.97. Additionally, video-level evaluation under realistic conditions yielded up to 0.79 accuracy and 0.84 AUC-PR. These results support the feasibility of integrating eye state detection into broader AI frameworks for neonatal monitoring and early neurological assessment.
This article presents a multivariate time series dataset detailing the physicochemical degradation of an industrial metalworking fluid (MWF). The data were collected continuously over several months from a test tank under typical operational conditions at an industrial facility in Spain. Four critical variables were monitored using industrial-grade sensors: pH, temperature, concentration, and conductivity. The dataset is provided in five CSV files. The primary file, measures.csv, contains the preprocessed time series at a uniform 5-minute frequency, with authentic missing data gaps intentionally preserved to reflect real-world sensor and connectivity issues. The four additional files serve as a comprehensive benchmark for data imputation algorithms. Each of these benchmark files corresponds to a single variable and includes the original data alongside imputed values generated by five distinct methods: K-Nearest Neighbours (KNN), a hybrid model (HybridKCL), an LSTM-based Variational Autoencoder (LSTM-VAE), and both pre-trained and fine-tuned versions of the MOMENT foundation model. This resource enables researchers and practitioners to develop, validate, and compare predictive maintenance models, anomaly detection systems, and advanced imputation techniques. Furthermore, it serves as a valuable educational tool for addressing common challenges in industrial IoT data, fostering advancements in sustainable and efficient manufacturing.
Background and objectiveVery preterm infants are highly susceptible to Neurodevelopmental Impairments (NDIs), including cognitive, motor, and language deficits. This paper presents a systematic review of the application of Machine Learning (ML) techniques to predict NDIs in premature infants.MethodsThis review presents a comparative analysis of existing studies from January 2018 to December 2023, highlighting their strengths, limitations, and future research directions.ResultsWe identified 26 studies that fulfilled the inclusion criteria. In addition, we explore the potential of ML algorithms and discuss commonly used data sources, including clinical and neuroimaging data. Furthermore, the inclusion of omics data as a contemporary approach employed, in other diagnostic contexts is proposed.ConclusionsWe identified limitations and emphasized the significance of employing multimodal data models and explored various alternatives to address the limitations identified in the reviewed studies. The insights derived from this review guide researchers and clinicians toward improving early identification and intervention strategies for NDIs in this vulnerable population.