Generating accurate SQL queries from natural language is critical for enabling non-experts to interact with complex databases, particularly in high-stakes domains like healthcare. This paper presents an extensive evaluation of state-of-the-art large language models (LLM), including LLaMA 3.3, Mixtral, Gemini, Claude 3.5, GPT-4o, and Qwen for transforming medical questions into executable SQL queries using the MIMIC-3 and TREQS datasets. Our approach employs LLMs with various prompts across 1000 natural language questions. The experiments are repeated multiple times to assess performance consistency, token efficiency, and cost-effectiveness. We explore the impact of prompt design on model accuracy through an ablation study, focusing on the role of table data samples and one-shot learning examples. The results highlight substantial trade-offs between accuracy, consistency, and computational cost between the models. This study also underscores the limitations of current models in handling medical terminology and provides insights to improve SQL query generation in the healthcare domain. Future directions include implementing RAG pipelines based on embeddings and reranking models, integrating ICD taxonomies, and refining evaluation metrics for medical query performance. By bridging these gaps, language models can become reliable tools for medical database interaction, enhancing accessibility and decision-making in clinical settings.
This paper presents a memetic algorithm (MA) for energy cost estimation of a robot path. The developed algorithm uses a random recombination genetic algorithm (GA) as the basis for the first stage of the algorithm and performs a local search based on feature importances determined from the data in the second stage. To allow for the faster determination of the solution quality, the algorithm uses an ML-driven fitness function, based on MLP, for the determination of path energy. The performed tests show that not only does the GA itself optimize the point-to-point paths well, but the usage of MA can lower the energy use by 58% on average (N = 100) when compared to a linear path between the same two points.
Large Language Models (LLMs) are increasingly recognized for their potential to alleviate administrative burdens in healthcare, enabling medical professionals to focus more on patient care rather than time-consuming administrative tasks. This paper explores how local LLMs can support healthcare settings by automating and streamlining routine administrative duties, improving workflow efficiency, and ultimately enhancing patient care.One of the key applications of LLMs is in the management of medical documentation. Healthcare professionals often spend significant time on tasks such as transcribing patient notes, updating medical records, and completing forms. By using LLMs, these processes can be automated or simplified. The models can transcribe and structure patient interactions in real-time, generate diagnostic summaries, and update electronic health records (EHRs) based on structured inputs, reducing the time healthcare providers spend on paperwork. This not only saves valuable time but also minimizes the risk of errors associated with manual data entry.Another area where LLMs can be beneficial is in appointment scheduling and patient communication. LLMs can be integrated into practice management systems to manage appointments, send reminders, and handle patient inquiries. By processing natural language requests, these models can schedule or reschedule appointments, direct patients to the appropriate specialists, and provide answers to frequently asked questions. This reduces the administrative workload for healthcare staff, allowing them to focus on more critical tasks and improving overall clinic efficiency.In addition, LLMs can assist with billing and insurance processing. By automatically extracting relevant information from patient records and claims, LLMs can generate billing codes, verify insurance coverage, and ensure that all necessary documentation is submitted. This reduces the administrative burden on healthcare providers and billing departments, streamlining the reimbursement process and minimizing errors in insurance claims.Local LLMs also aid in regulatory compliance by automatically ensuring that healthcare institutions adhere to relevant legal requirements, such as patient consent and privacy regulations. By continuously monitoring the creation and modification of medical records, these models can flag potential issues related to compliance and generate alerts, ensuring that healthcare providers remain in line with regulations such as GDPR. This proactive approach to compliance reduces the risk of legal liabilities and minimizes the time spent on manual checks.In conclusion, by automating tasks such as documentation, scheduling, billing, and regulatory compliance, local LLMs can improve workflow efficiency, reduce human error, and enhance overall productivity in healthcare settings. As the technology evolves, local LLMs have the potential to significantly transform healthcare operations, allowing medical professionals to focus on what matters most: their patients.
This paper focuses on the estimation of electrical power output (Pe) in a combined cycle power plant (CCPP) using ambient temperature (AT), vacuum in the condenser (V), ambient pressure (AP), and relative humidity (RH). The study stresses accurate estimation for better CCPP performance and energy efficiency through responsive control to changing conditions. The novelty lies in applying genetic programming (GP) on a publicly available dataset to generate Symbolic Expressions (SEs) for high-accuracy Pe. To address the challenge of numerous GP hyperparameters, a random hyperparameter values search method (RHVS) is introduced to find optimal combinations, resulting in SEs with higher accuracy. SEs are created with varying input variables, and their performance is evaluated using multiple metrics (coefficient of determination (R2), mean absolute error (MAE), mean square error (MSE), root mean square error (RMSE), mean absolute percentage error (MAPE), Kling–Gupta Efficiency (KGE), and Bland–Altman (B-A) analysis). A key innovation involves combining the best SEs through an Averaging ensemble (AE), leading to a robust estimation accuracy. Notably, the AE YVE−2 achieves the highest (Pe) accuracy, including R2=0.9368, MAE=3.3378, MSE=18.4800, RMSE=4.2985, MAPE=0.7354%, and KGE=0.9479. The investigation highlights AT as the most influential variable, underscoring the importance of choosing inputs aligned with physical processes. This paper’s outlined procedure, combining GP, hyperparameter optimization, and ensemble techniques, offers an efficient method for estimating Pe in CCPP. It promises simplicity and effectiveness in real-world applications. B-A analysis proves valuable for SE selection, enhancing the proposed methodology.
Maritime security and monitoring are essential for global trade, environmental protection, and national defense. Traditional machine learning models have been effective in recognizing and classifying maritime objects, but their reliance on large, labeled datasets poses significant challenges, particularly in dynamic environments where new and unforeseen objects frequently emerge. This study explores the application of Zero-Shot Learning (ZSL) to the maritime domain, leveraging the CLIP (Contrastive Language-Image Pre-training) model to classify maritime objects with minimal labeled data. A custom dataset comprising 1,438 images across five maritime object categories- boat, cargo, cruise, dock, and lighthouse-was curated for evaluation. Four CLIP model variants were examined: "clip-vit-base-patch16," "clip-vit-base-patch32," "clip-vit-large-patch14," and "clipvit-large-patch14-336." The study's findings indicate that the CLIP models, particularly the "clip-vitlarge-patch14-336" variant, achieve high classification accuracy, with AUC values approaching 1.0 across most classes. Performance was strongest in easily distinguishable categories such as dock and lighthouse, but challenges remain with rare or ambiguous classes such as cargo ships, where F2 scores suggest variability in recall and precision. Additionally, the study highlights the potential limitations of these models, including their dependency on dataset diversity and potential biases introduced by web-scraped images, which may not fully represent the complex, real-world conditions of maritime environments.
Maritime security and monitoring are essential for global trade, environmental protection,and national defense. Traditional machine learning models have been effective in recognizing andclassifying maritime objects, but their reliance on large, labeled datasets poses challenges,particularly in dynamic environments where new and unforeseen objects frequently emerge. Thisstudy explores the application of Zero-Shot Learning (ZSL) to the maritime domain, leveraging theCLIP model to classify maritime objects with minimal labeled data. A custom dataset comprising1,438 images was used to evaluate the performance of various CLIP model variants. Our findingsindicate that CLIP models, particularly the "clip-vit-large-patch14-336" variant, achieve highclassification accuracy, with AUC values approaching 1.0 across most classes. However, challengesremain in handling rare or ambiguous classes such as cargo ships, where F2 scores suggestvariability in recall and precision. Additionally, the study highlights the potential limitations ofthese models, including their dependency on dataset diversity and the risk of overfitting to specificdata characteristics. The "clip-vit-large-patch14-336" model is identified as the most balanced andreliable option, offering a strong foundation for enhancing maritime situational awareness andsupporting diverse maritime applications.
Predicting the parameters of a Combined Diesel-Electric and Gas (CODLAG) propulsion system is crucial for optimizing the design, performance, and reliability of these complex engineering systems. This paper presents the implementation of the Kolmogorov-Arnold Network (KAN) for predicting CODLAG system parameters. The KAN, based on Kolmogorov's superposition theorem, can accurately approximate nonlinear relationships within the system. We provide a comprehensive overview of CODLAG systems, their operational principles, and the challenges in parameter prediction. Experimental results demonstrate that increasing hidden layers, optimizing learning rates, and adjusting batch sizes and epochs significantly improve prediction accuracy. The findings highlight the robustness of KAN in modeling complex systems and pave the way for future research to enhance the predictive capabilities of engineering systems.
This study delves into the vital missions of the armed forces, encompassing the defense of territorial integrity, sovereignty, and support for civil institutions. Commanders grapple with crucial decisions, where accountability underscores the imperative for reliable field intelligence. Harnessing artificial intelligence, specifically, the YOLO version five detection algorithm, ensures a paradigm of efficiency and precision. The presentation of trained models, accompanied by pertinent hyperparameters and dataset specifics derived from public military insignia videos and photos, reveals a nuanced evaluation. Results scrutinized through precision, recall, map@0.5, mAP@0.95, and F1 score metrics, illuminate the supremacy of the model employing Stochastic Gradient Descent at 640 × 640 resolution: 0.966, 0.957, 0.979, 0.830, and 0.961. Conversely, the suboptimal performance of the model using the Adam optimizer registers metrics of 0.818, 0.762, 0.785, 0.430, and 0.789. These outcomes underscore the model’s potential for military object detection across diverse terrains, with future prospects considering the implementation on unmanned arial vehicles to amplify and deploy the model effectively.
Neurological diseases pose a significant public health challenge, leading to disability and mortality globally.Current diagnostics for neuroinflammatory diseases are complex and lack efficacy, necessitating invasive procedures.Calcium signaling dynamics in astrocytes and microglia play pivotal roles in central nervous system (CNS) function and dysfunction.This study curates time-series data on calcium transients in astrocytes and microglia, employing live-cell imaging techniques and preprocessing methodologies.Using k-means clustering, we analyze the data, revealing optimal clustering solutions between 2 to 6 clusters.This research offers valuable insights for understanding CNS disorders and highlights the potential of clustering techniques in neurological research.
The development of artificial intelligence is one of the most significant technological innovations that contributes to humanity with its characteristics and facilitates, secures, and improves everyday life. However the challenge arises when the programmer-engineer learns the algorithm of artificial intelligence to perform the proposed task, that is, the challenge lies in the issue of available hardware resources. Artificial intelligence algorithms perform many human-impossible tasks, such as detection and counting, i.e. calculating the interrelationships of individual objects, segmentation of tumors and other malignant diseases, i.e. tissues, classification of specific states of classes of a scene, and many other similar technologies. This research paper examines the possibility of implementing the You Only Look Once algorithm of the tenth generation on certain devices such as an affordable Raspberry Pi, and will discuss the advantages and disadvantages of changes in detection parameters, i.e. inferences to the applied model. In addition, a mock-up of the device will be shown, which will serve to provide timely information about criminals and suspicious persons who possess firearms in different situations, such as normal weather conditions in a populated place or in shops where petty robberies are frequent. The testing will be done using recorded videos in real-time scenarios. Finally, real-time inference or detection in real-time will be tested and the actions that Raspberry will perform will be simulated. The results indicate that the optimal model achieved a precision of 0.938, recall of 9.863, mAP50 of 0.91, and mAP50-95 of 0.739. This was achieved using an image size of 640, IOU threshold of 0.7, confidence threshold of 0.6, and training for 600 iterations with the Stochastic Gradient Descent optimizer, without augmentations, and employing the ONNX inference format.
Individual safety in urban and city centers is one of the fundamental rights of every person. Today's innovations open up various possibilities for improving the quality of life and safety, but if they are not applied, they carry the danger of an additional increase in crime or greater insecurity in urban centers. This investigation aims to apply insights into the anticipation of illegal actions by various criminal groups that use illegal means, such as weapons, to get property that does not belong to them. In this research, computer vision and advanced strategies such as YOLOv10 and YOLOv9 are used to compare their properties and performance for improving security in urban and populated areas. The best results were achieved by YOLOv9m with a high percentage of mean average precision in the amount of 0.923, while the worst performance was shown by YOLOv10n with a mean average precision of 0.902. Precision and recall for the YOLOv9 and YOLOv10 models were also taken into account and the results are as follows: 0.984.. 0.993.. 0.979.. and 0.996.
A synchronous machine is an electro-mechanical converter consisting of a stator and a rotor. The stator is the stationary part of a synchronous machine that is made of phase-shifted armature windings in which voltage is generated and the rotor is the rotating part made using permanent magnets or electromagnets. The excitation current is a significant parameter of the synchronous machine, and it is of immense importance to continuously monitor possible value changes to ensure the smooth and high-quality operation of the synchronous machine itself. The purpose of this paper is to estimate the excitation current on a publicly available dataset, using the following input parameters: Iy: load current; PF: power factor; e: power factor error; and df: changing of excitation current of synchronous machine, using artificial intelligence algorithms. The algorithms used in this research were: k-nearest neighbors, linear, random forest, ridge, stochastic gradient descent, support vector regressor, multi-layer perceptron, and extreme gradient boost regressor, where the worst result was elasticnet, with R2 = −0.0001, MSE = 0.0297, and MAPE = 0.1442; the best results were provided by extreme boosting regressor, with R2¯ = 0.9963, MSE¯ = 0.0001, and MAPE¯ = 0.0057, respectively.
The navigation of mobile robots throughout the surrounding environment without collisions is one of the mandatory behaviors in the field of mobile robotics. The movement of the robot through its surrounding environment is achieved using sensors and a control system. The application of artificial intelligence could potentially predict the possible movement of a mobile robot if a robot encounters potential obstacles. The data used in this paper is obtained from a wall-following robot that navigates through the room following the wall in a clockwise direction with the use of 24 ultrasound sensors. The idea of this paper is to apply genetic programming symbolic classifier (GPSC) with random hyperparameter search and 5-fold cross-validation to investigate if these methods could classify the movement in the correct category (move forward, slight right turn, sharp right turn, and slight left turn) with high accuracy. Since the original dataset is imbalanced, oversampling methods (ADASYN, SMOTE, and BorderlineSMOTE) were applied to achieve the balance between class samples. These over-sampled dataset variations were used to train the GPSC algorithm with a random hyperparameter search and 5-fold cross-validation. The mean and standard deviation of accuracy (ACC), the area under the receiver operating characteristic (AUC), precision, recall, and F1−score values were used to measure the classification performance of the obtained symbolic expressions. The investigation showed that the best symbolic expressions were obtained on a dataset balanced with the BorderlineSMOTE method with ACC¯±SD(ACC), AUC¯macro±SD(AUC), Precision¯macro±SD(Precision), Recall¯macro±SD(Recall), and F1−score¯macro±SD(F1−score) equal to 0.975×1.81×10−3, 0.997±6.37×10−4, 0.975±1.82×10−3, 0.976±1.59×10−3, and 0.9785±1.74×10−3, respectively. The final test was to use the set of best symbolic expressions and apply them to the original dataset. In this case the ACC¯±SD(ACC), AUC¯±SD(AUC), Precision¯±SD(Precision), Recall¯±SD(Recall), and F1−score¯±SD(F1−Score) are equal to 0.956±0.05, 0.9536±0.057, 0.9507±0.0275, 0.9809±0.01, 0.9698±0.00725, respectively. The results of the investigation showed that this simple, non-linearly separable classification task could be solved using the GPSC algorithm with high accuracy.
Hepatitis C is an infectious disease which is caused by the Hepatitis C virus (HCV) and the virus primarily affects the liver. Based on the publicly available dataset used in this paper the idea is to develop a mathematical equation that could be used to detect HCV patients with high accuracy based on the enzymes, proteins, and biomarker values contained in a patient’s blood sample using genetic programming symbolic classification (GPSC) algorithm. Not only that, but the idea was also to obtain a mathematical equation that could detect the progress of the disease i.e., Hepatitis C, Fibrosis, and Cirrhosis using the GPSC algorithm. Since the original dataset was imbalanced (a large number of healthy patients versus a small number of Hepatitis C/Fibrosis/Cirrhosis patients) the dataset was balanced using random oversampling, SMOTE, ADSYN, and Borderline SMOTE methods. The symbolic expressions (mathematical equations) were obtained using the GPSC algorithm using a rigorous process of 5-fold cross-validation with a random hyperparameter search method which had to be developed for this problem. To evaluate each symbolic expression generated with GPSC the mean and standard deviation values of accuracy (ACC), the area under the receiver operating characteristic curve (AUC), precision, recall, and F1-score were obtained. In a simple binary case (healthy vs. Hepatitis C patients) the best case was achieved with a dataset balanced with the Borderline SMOTE method. The results are ACC¯±SD(ACC), AUC¯±SD(AUC), Precision¯±SD(Precision), Recall¯±SD(Recall), and F1−score¯±SD(F1−score) equal to 0.99±5.8×10−3, 0.99±5.4×10−3, 0.998±1.3×10−3, 0.98±1.19×10−3, and 0.99±5.39×10−3, respectively. For the multiclass problem, OneVsRestClassifer was used in combination with GPSC 5-fold cross-validation and random hyperparameter search, and the best case was achieved with a dataset balanced with the Borderline SMOTE method. To evaluate symbolic expressions obtained in this case previous evaluation metric methods were used however for AUC, Precision, Recall, and F1−score the macro values were computed since this method calculates metrics for each label, and find their unweighted mean value. In multiclass case the ACC¯±SD(ACC), AUC¯macro±SD(AUC), Precision¯macro±SD(Precision), Recall¯macro±SD(Recall), and F1−score¯macro±SD(F1−score) are equal to 0.934±9×10−3, 0.987±1.8×10−3, 0.942±6.9×10−3, 0.934±7.84×10−3 and 0.932±8.4×10−3, respectively. For the best binary and multi-class cases, the symbolic expressions are shown and evaluated on the original dataset.
In the case of pandemics such as COVID-19, the rapid development of medicines addressing the symptoms is necessary to alleviate the pressure on the medical system. One of the key steps in medicine evaluation is the determination of pIC50 factor, which is a negative logarithmic expression of the half maximal inhibitory concentration (IC50). Determining this value can be a lengthy and complicated process. A tool allowing for a quick approximation of pIC50 based on the molecular makeup of medicine could be valuable. In this paper, the creation of the artificial intelligence (AI)-based model is performed using a publicly available dataset of molecules and their pIC50 values. The modeling algorithms used are artificial and convolutional neural networks (ANN and CNN). Three approaches are tested-modeling using just molecular properties (MP), encoded SMILES representation of the molecule, and the combination of both input types. Models are evaluated using the coefficient of determination (R2) and mean absolute percentage error (MAPE) in a five-fold cross-validation scheme to assure the validity of the results. The obtained models show that the highest quality regression (R2¯=0.99, σR2¯=0.001; MAPE¯=0.009%, σMAPE¯=0.009), by a large margin, is obtained when using a hybrid neural network trained with both MP and SMILES.
Printed circuit boards (PCBs) are an indispensable part of every electronic device used today. With its computing power, it performs tasks in much smaller dimensions, but the process of making and sorting PCBs can be a challenge in PCB factories. One of the main challenges in factories that use robotic manipulators for “pick and place” tasks are object orientation because the robotic manipulator can misread the orientation of the object and thereby grasp it incorrectly, and for this reason, object segmentation is the ideal solution for the given problem. In this research, the performance, memory size, and prediction of the YOLO version 5 (YOLOv5) semantic segmentation algorithm are tested for the needs of detection, classification, and segmentation of PCB microcontrollers. YOLOv5 was trained on 13 classes of PCB images from a publicly available dataset that was modified and consists of 1300 images. The training was performed using different structures of YOLOv5 neural networks, while nano, small, medium, and large neural networks were used to select the optimal network for the given challenge. Additionally, the total dataset was cross validated using 5-fold cross validation and evaluated using mean average precision, precision, recall, and F1-score classification metrics. The results showed that large, computationally demanding neural networks are not required for the given challenge, as demonstrated by the YOLOv5 small model with the obtained mAP, precision, recall, and F1-score in the amounts of 0.994, 0.996, 0.995, and 0.996, respectively. Based on the obtained evaluation metrics and prediction results, the obtained model can be implemented in factories for PCB sorting applications.
Objectives: Cervical cancer is present in most cases of squamous cell carcinoma. In most cases, it is the result of an infection with human papillomavirus or adenocarcinoma. This type of cancer is the third most common cancer of the female reproductive organs. The risk groups for cervical cancer are mostly younger women who frequently change partners, have early sexual intercourse, are infected with human papillomavirus (HPV), and who are nicotine addicts. In most cases, the cancer is asymptomatic until it has progressed to the later stages. Cervical cancer screening rates are low, especially in developing countries and in some minority groups. Due to these facts, the introduction of a tentative cervical cancer screening based on a questionnaire can enable more diagnoses of cervical cancer in the initial stages of the disease. Methods: In this research, publicly available cervical cancer data collected on 859 female patients are used. Each sample consists of 36 input attributes and four different outputs Hinselmann, Schiller, cytology, and biopsy. Due to the significant unbalance of the data set, class balancing techniques were used, and these are the Synthetic Minority Oversampling Technique, the ADAptive SYNthetic algorithm (ADASYN), SMOTEEN, random oversampling, and SMOTETOMEK. To obtain the mentioned target outputs, multiple artificial intelligence (AI) and machine learning (ML) methods are proposed. In this research, multiple classification algorithms such as logistic regression, multilayer perceptron (MLP), support vector machine (SVM), K-nearest neighbors (KNN), and several naive Bayes methods were used. Results: From the achieved results, it can be seen that the highest performances were achieved if MLP and KNN are used in combination with Random oversampling, SMOTEEN, and SMOTETOMEK. Such an approach has resulted in mean area under the receiver operating characteristic curve (AUC¯) and mean Matthew’s correlation coefficient (MCC¯) scores of higher than 0.95, regardless of which diagnostic method was used for output vector construction. Conclusions: According to the presented results, it can be concluded that there is a possibility for the utilization of artificial intelligence (AI) and machine learning (ML) techniques for the development of a tentative cervical cancer screening method, which is based on a questionnaire and an AI-based algorithm. Furthermore, it can be concluded that by using class balancing techniques, a certain performance boost can be achieved.
Path planning is one of the key steps in the application of industrial robotic manipulators. The process of determining trajectories can be time-intensive and mathematically complex, which raises the complexity and error proneness of this task. For these reasons, the authors tested the application of a genetic algorithm (GA) on the problem of continuous path planning based on the Ho–Cook method. The generation of trajectories was optimized with regard to the distance between individual segments. A boundary condition was set regarding the minimal values that the trajectory parameters can be set in order to avoid stationary solutions. Any distances between segments introduced by this condition were addressed with Bezier spline interpolation applied between evolved segments. The developed algorithm was shown to generate trajectories and can easily be applied for the further path planning of various robotic manipulators, which indicates great promise for the use of such algorithms.
Imaging is one of the main tools of modern astronomy-many images are collected each day, and they must be processed. Processing such a large amount of images can be complex, time-consuming, and may require advanced tools. One of the techniques that may be employed is artificial intelligence (AI)-based image detection and classification. In this paper, the research is focused on developing such a system for the problem of the Magellan dataset, which contains 134 satellite images of Venus's surface with individual volcanoes marked with circular labels. Volcanoes are classified into four classes depending on their features. In this paper, the authors apply the You-Only-Look-Once (YOLO) algorithm, which is based on a convolutional neural network (CNN). To apply this technique, the original labels are first converted into a suitable YOLO format. Then, due to the relatively small number of images in the dataset, deterministic augmentation techniques are applied. Hyperparameters of the YOLO network are tuned to achieve the best results, which are evaluated as mean average precision (mAP@0.5) for localization accuracy and F1 score for classification accuracy. The experimental results using cross-vallidation indicate that the proposed method achieved 0.835 mAP@0.5 and 0.826 F1 scores, respectively.
Fire is usually detected with fire detection systems that are used to sense one or more products resulting from the fire such as smoke, heat, infrared, ultraviolet light radiation, or gas. Smoke detectors are mostly used in residential areas while fire alarm systems (heat, smoke, flame, and fire gas detectors) are used in commercial, industrial and municipal areas. However, in addition to smoke, heat, infrared, ultraviolet light radiation, or gas, other parameters could indicate a fire, such as air temperature, air pressure, and humidity, among others. Collecting these parameters requires the development of a sensor fusion system. However, with such a system, it is necessary to develop a simple system based on artificial intelligence (AI) that will be able to detect fire with high accuracy using the information collected from the sensor fusion system. The novelty of this paper is to show the procedure of how a simple AI system can be created in form of symbolic expression obtained with a genetic programming symbolic classifier (GPSC) algorithm and can be used as an additional tool to detect fire with high classification accuracy. Since the investigation is based on an initially imbalanced and publicly available dataset (high number of samples classified as 1-Fire Alarm and small number of samples 0-No Fire Alarm), the idea is to implement various balancing methods such as random undersampling/oversampling, Near Miss-1, ADASYN, SMOTE, and Borderline SMOTE. The obtained balanced datasets were used in GPSC with random hyperparameter search combined with 5-fold cross-validation to obtain symbolic expressions that could detect fire with high classification accuracy. For this investigation, the random hyperparameter search method and 5-fold cross-validation had to be developed. Each obtained symbolic expression was evaluated on train and test datasets to obtain mean and standard deviation values of accuracy (ACC), area under the receiver operating characteristic curve (AUC), precision, recall, and F1-score. Based on the conducted investigation, the highest classification metric values were achieved in the case of the dataset balanced with SMOTE method. The obtained values of ACC¯±SD(ACC), AUC¯±SD(ACU), Precision¯±SD(Precision), Recall¯±SD(Recall), and F1-score¯±SD(F1-score) are equal to 0.998±4.79×10−5, 0.998±4.79×10−5, 0.999±5.32×10−5, 0.998±4.26×10−5, and 0.998±4.796×10−5, respectively. The symbolic expression using which best values of classification metrics were achieved is shown, and the final evaluation was performed on the original dataset.