
The application of artificial intelligence (AI) in civil engineering presents a transformative approach to enhancing design quality and safety. This paper investigates the potential of the advanced LLM GPT4 Turbo vision model in detecting architectural flaws during the design phase, with a specific focus on identifying missing doors and windows. The study evaluates the model's performance through metrics such as precision, recall, and F1 score, demonstrating AI's effectiveness in accurately detecting flaws compared to human-verified data. Additionally, the research explores AI's broader capabilities, including identifying load-bearing issues, material weaknesses, and ensuring compliance with building codes. The findings highlight how AI can significantly improve design accuracy, reduce costly revisions, and support sustainable practices, ultimately revolutionizing the civil engineering field by ensuring safer, more efficient, and aesthetically optimized structures.
The integration of Unmanned Aerial Vehicles (UAVs) into existing cloud-based fire safety systems provides a transformative approach to fire detection, monitoring, and emergency response. This paper proposes an extension of the existing ALIKE system to incorporate UAV-based 3D mapping and hotspot detection capabilities. The enhanced system aims to provide real-time situational awareness, reduce response times, and improve the safety and effectiveness of firefighting operations. The implementation of UAV technology, combined with cloud computing and AI, represents a significant advancement in fire safety management. Utilizing UAVs in fire safety offers benefits such as high-resolution aerial imagery, thermal imaging for precise hotspot detection, and the ability to access hazardous areas without endangering human lives. This integration facilitates continuous monitoring and data collection, enabling predictive analytics for fire spread and risk assessment. By leveraging the strengths of UAVs, cloud computing, and AI, the proposed system aims to enhance the overall resilience and responsiveness of fire safety operations.
Pedestrian safety is a topic that has been tackled by many researchers in the past using machine learning and artificial intelligence approaches. One of the main reasons for this is the development of self-driving vehicles, which would require systems to detect pedestrians to ensure their safety. Many approaches utilize onboard camera systems which work great for most situations but there are certain shortcomings to this approach. One case is with blind right turns where the onboard cameras cannot see the entire crosswalk due to obstruction of a parked vehicle or building close to the corner. An alternative approach is to use cameras that already exist at intersections to create a warning system that detects pedestrians at each crosswalk. There are many methods of implementation, such as integration with self-driving vehicles using a vehicle-to-infrastructure (V2I) approach or using a caution light to warn drivers in advance. This paper will discuss approaches and considerations for using object detection for intersection safety using cameras at intersections. Results using different pre-trained YOLOv8 models provided by Ultralytics for three different intersection datasets from Urban Tracker: Multiple Object Tracking in Urban Mixed Traffic show that the YOLOv8 pre-trained models perform extremely well for this task.
Autism Spectrum Disorder (ASD) is a neurological and developmental condition that presents considerable social, behavioral, and communicative challenges to those diagnosed with it. Recent advances in music therapy have shown promising results in addressing these challenges. This review examines studies on music therapy for ASD from 2020 to 2024, covering both biomedical and interdisciplinary repositories. Our analysis reveals that music therapy can effectively enhance communication and social skills in people with ASD. Based on the results from 18 selected studies, the most significant improvements are achieved in social-emotional reciprocity and adaptive behaviors, contrasting with only modest improvements in social relationship deficits. Among the intervention approaches assessed, improvising and playing an instrument achieved slightly better outcomes. However, while promising, the robustness of such conclusions is bound by small sample sizes, insufficient longitudinal studies, and limited age diversity among participants. These limitations underscore the need for largerscale, more diverse, and more in-depth research to confirm the findings and extend their generalizability. Previous reviews have discussed the application of technologies such as robots, virtual reality (VR), and wearable devices in music therapy for ASD. This review expands on these discussions by exploring the integration of artificial intelligence (AI), which offers significant potential to enhance the personalization, adaptiveness, and efficacy of traditional music interventions for ASD. We explore promising avenues for multifaceted AI applications in music selection and generation, and automatic emotion and response detection, while also discussing the challenges of implementing AI in music therapy for ASD.
Data exhibit the distribution of the problem space, and the efficacy of machine learning models is contingent upon the availability of quality datasets. Additionally, in traditional machine learning models, data are required to be collected on a centralized server for training and testing, which raises privacy and security concerns. Federated Learning (FL) was designed with the aim of moving the computation to the data source (clients) instead of bringing the data to the central server. While addressing data accessibility and privacy issues, the eligibility of clients and the quality of their data should be considered in a secure FL framework. It is difficult to verify the authenticity of the clients participating in the FL environment. Various solutions have been proposed, including statistical analysis of client updates, hardware-based isolation, Differential Privacy (DP), Homomorphic Encryption (HE), and others. However, these solutions still pose limitations and suffer significantly from trade-offs, such as the privacy-utility tradeoff. In this research, we propose an approach to fortify the FL environment with continuous verification of client updates to prevent model poisoning attacks and use a filter ensemble to detect the data poisoning attacks. Our empirical experiments demonstrated improved performance against these attacks and alleviated the limitations in existing solutions.
Critical infrastructures, such as water treatment plants (WTPs) and communication networks, are vital to our daily lives, providing essential resources. As these infrastructures increasingly rely on computer systems and internet networks, conducting cybersecurity training on expensive, real-world equipment becomes impractical due to the risk of operational interruptions. To address this challenge, we present a Digital Twin training platform that simulates three key cybersecurity concepts: Input Manipulation, Output Manipulation, and Denial of Service attacks, within the context of WTPs. These concepts are mapped to four critical functions: water level, chlorine level, water temperature, and the microbial water purification process. This paper primarily focuses on the design and development of the VR component of our digital twin platform, as well as the integration process with our hardware testbed. Initial investigations demonstrated significant potential for the experiential learning platform, serving as an effective tool for educating users about cybersecurity issues in mission-critical facilities such as WTPs.
As Large Language Models (LLMs) continue to advance and be utilized across societal domains, such as finance, healthcare, and the justice system, the need to address the inherent bias and the ability to learn new biases in these models becomes imminent. Concerns regarding these biases are rising, given their potential to perpetuate and even amplify existing social inequalities. This paper explores the multifaceted nature of bias in artificial intelligence (AI), examining the similarities and differences between human and machine bias. We delve into the origins of bias, distinguishing between those introduced by users and those inherent in the AI systems themselves. Our study focuses on the mechanisms by which biases are elicited and amplified through human-to-machine interactions. Through experimentation and analysis, we implement methodologies for eliciting, measuring, and mitigating these biases. Our results suggest that even though LLMs like ChatGPT-4 are equipped with effective content moderators, these chatbots can still learn and exhibit biased responses through human coercion. Further, we have learned that these biases are both inherent and learned through human interaction. Finally, we offer insightful strategies to mitigate these biases in LLMs.
Deepfake technology’s rise has led to a surge in false identities, creating a significant and present problem with broad societal ramifications. Concerns over identity theft, harassment, and the dissemination of false information have escalated due to the simplicity with which deepfaked facial images can now be produced and distributed thanks to the broad availability of generative AI tools like Generative Adversarial Networks (GANs). The availability of these tools has political ramifications since it can degrade public opinion and damage institutional trust. As such, the ability to identify deepfake face images has become essential. Ensuring a person’s identity is critical in preventing the dissemination of false information on social media. Detection of deepfake facial images is also necessary for identity verification in border control, law enforcement, and security applications. To effectively and precisely recognize deepfake face images, this study effort has focused on modifying transfer learning models, such as ResNet101V2, MobileNetV2, NASNetLarge, NASNetMobile, DenseNet121, DenseNet169, DenseNet201, and Xception.
In this project, a design for an automated mechanism is presented for implementation in textile manufacturing processes, where an integrated system for textile fiber recognition based on artificial intelligence (AI) was added. The mechanical design of the machine was studied in terms of force, providing a robust and efficient structure that supports the integration of sensors and cameras for real-time image capture of the processed fibers. Additionally, this design included a precise feeding system and a transport mechanism that ensured the stability and correct positioning of the fibers during analysis. In parallel, an AI model was implemented to identify and classify textile fibers based on their color characteristics. A simulated dataset was generated, where each type of fiber was represented by a specific color (red, green, blue, yellow), and a simple neural network was trained to recognize these color patterns. The model was optimized to achieve high accuracy in fiber classification and was subsequently evaluated with a test set. The system also included a visualization functionality that allowed the recognized fiber color to be displayed along with its classification, providing visual validation of the process. This comprehensive approach, combining advanced mechanical design with AI, proved effective in improving the accuracy and efficiency of automatic textile fiber identification, significantly contributing to the optimization of production processes in the textile industry.
The global appeal and competitive nature of soccer have led to the application of technology to monitor and enhance player training and performance. Effective and data-driven training can reduce injury related to over-exertion and boost player performance. Shooting is a crucial skill in soccer and can significantly influence the outcome of a game. This work addresses the lack of a smart and privacy-aware device to track and analyze players' shot localization accuracy, particularly for amateur players. We investigate a smart soccer net that can localize the point of impact of the ball from the perturbations of small inertial measuring sensors attached to the net. We train a feedforward neural network with raw data and statistical-based features, and a convolutional neural network with time- and frequency-based features from continuous wavelet transformation (CWT). Also, we investigated one, two, three, and four sensors deployed in selected configurations on the net. Four sensors deployed in a rhombus configuration and a model trained with CWT-based features achieved the best classification accuracy of 94%. A cost-effective option of using three sensors in an inverted triangle configuration and the model trained with statistical-based features achieved 90% accuracy. This work is a step towards a minimalistic system that can be used as a training aid to automatically track the positions of soccer player shots in the goal during training sessions.
Eye-tracking technology has long been a cornerstone in both academic and research fields, offering insights into behavior, cognition, and visual perception. That being said, however, its accessibility is hindered by the high costs and proprietary nature of existing methodologies. To address these issues, we present AITracker, an open-source application that works using only a standard webcam, leveraging a deep-learning model in order to provide accurate eye-tracking capabilities. Our solution offers flexibility and adaptability, enabling users to customize individual parameters in the software to meet their specific needs. Through robust data collection and neural network training, AITracker achieves incredibly fast response times with a high degree of accuracy, enabling gaze-tracking in up to eight directions, as well as blink detection. In order to better understand the impact of this technology in the context of existing solutions, this paper compares AITracker’s multi-layered neural network to various other prevalent eye-tracking methodologies. To that end, this paper also notes certain limitations that inhibit the software, including an undersized dataset and problematic distribution. Additionally, we explore various application scenarios, including hardware integration for assistive technology, hands-free gaming interfaces, advertising research, and attention monitoring in education. Moreover, feedback gathered from different users highlights the effectiveness and impact of AITracker across a diverse array of contexts.
A data-driven bus efficiency prediction model is a useful tool for transport planners to optimize the current and plan efficient new routes. This study proposes a novel approach for predicting efficiency scores by leveraging non radial DEA model and machine learning (ML) techniques. A labeled dataset is developed using a non-radial DEA method that considers interrelationships between operational and service efficiency and the selected features. Two machine learning models, Linear Regression (LR) and Support Vector Regression (SVR) are trained on the labeled dataset. The trained model can be used to predict the efficiency scores of a new bus routes based on decision-makers’ preferences on input parameters and without requiring a full DEA analysis. The methodology is experimented on CyRide, a real-world dataset provided by the Ames transit Agency. The effectiveness of both models in predicting efficiency is also evaluated using $R^{2}$, MSE, and residual plots with detailed discussion on exploration analysis of selected features and overall efficiency score. The proposed methodology can be generalized on any bus route dataset and used by transportation authorities for improved decision-making.
Reducing the utilization of energy is a core strategy for many goals related to environmental sustainability. Understanding which electrical appliances were deployed and how much power was consumed can be a first step toward identifying areas for greater efficiency in electrical usage. Non-intrusive load monitoring (NILM) is a technique for identifying the kinds of electrical appliances that are utilized based only on the measurable power readings. This study develops an improved framework for predictive modeling in NILM applications using the SustDataED2 household appliance data. We propose focusing the analysis only on activated periods of power consumption for each appliance and engineering features that are both predictive and simple to interpret. In the SustDataED2 records, the periods of activation collectively encompass approximately 99% of the overall power consumption but only account for 18.1% of the measured records. Within each period of activation, we calculate the following features: duration of activation, peak power consumption, mean power consumption, time to peak power, and the number of local peaks before the maximum value is reached. We also record the first five power readings for the period of activation. Using these features, we construct machine learning classification models in retrospective lookback models and as real-time classifiers. The lookback random forest model achieved a classification accuracy of 97.8%. The real-time random forest models ranged in accuracy from 84.9% after 1 observation (2 seconds) to 96.8% accuracy after 5 observations (10 seconds). We further demonstrate that the accuracy of these models is mainly limited by small sample sizes in some classes of appliances. These results improve upon prior research in both their accuracy and their ease of interpretation.
Recent advancements in artificial intelligence emphasize improving the abstraction capabilities of Large Language Models (LLMs) by integrating structured knowledge and sophisticated neural architectures. This paper proposes a novel approach grounded in a fourfold pattern observed at multiple levels of granularity in nature, referred to as 'Kn', 'Po', 'Pr', and 'H', which correlates with concepts of knowledge, presence, power, and harmony. We hypothesize that leveraging this intrinsic structure in language could enhance LLMs' ability to abstract underlying themes, recognize subtle relationships, and infer unstated implications, similar to the hierarchical complexity seen in natural systems from quantum particles to cellular structures. By incorporating these elements into the transformer architecture of LLMs, by enhancing self-attention through additional Q, K, V weights to thereby also optimize computational requirements, we aim to demonstrate a more nuanced understanding and generation of text, at least paralleling and perhaps even moving beyond normal human-like comprehension and reasoning.
Broadband high-altitude platform station (HAPS) concepts have garnered notable attention from 6G researchers due to their superior line-of-sight (LoS) probabilities, making them ideal for millimeter-wave (mmWave) and higher RF carrier frequencies. However, beam management processes for mmWave and higher carrier frequencies introduce substantial overhead. This paper analyzes vision-assisted HAPS 6G wireless communications using mmWave signaling. Machine Learning (ML) methods generate vision-based contextual information to facilitate mmWave beam selection, significantly reducing the need for extensive channel state information. The paper also analyzes a framework for vision-assisted millimeter-wave beam management that leverages visual information obtained by long-range vision sensors at the HAPS-gNB. Deep reinforcement learning methods, a highly efficient solution, provide intelligent assistance for managing 6G HAPS communications. The 6G Vision-Intelligent (6G-VI) ML system analysis utilizes a YOLOv8 deep learning model to identify the locations of User Equipment (UEs), which are then used as context for a millimeter-wave (mmWave) beam selection algorithm. The paper analyzes the 6G-VI method compared to other ML-based models, including Cascade Mask R-CNN, Mask R-CNN, Faster R-CNN, RetinaNet, and non-vision models. Simulation results indicate a minimum spectral efficiency improvement of 5 bits/sec/Hz compared to other prior aerial HAPS systems.
In this paper, we introduce a person-tracking method that utilizes a metric-based learning approach for object association during tracking. We evaluated the performance of this person association algorithm using various state-of-the-art detectors like Yolo V4, Resent, Faster RCNN, etc. based on the MOT dataset. Our main objective was to improve the accuracy and precision of multiple object tracking and to reduce ID switching, which are some of the problems related to various object tracking methods. Our approach leverages powerful detectors to effectively handle a large number of detected objects. The use of accurate feature extraction provides the metric learning methods with the opportunity to make correct associations with the right identity. This method is also capable of solving the person re-identification problem.
Academic emergency physicians and data scientists collaborated on a research program to reduce false alarms and to increase the utility of patient monitor systems for detecting significant cardiopulmonary events and clinical deterioration. An experimental hardware-software framework to study patient monitoring and alarm fatigue mitigation was implemented in 15 Emergency Department (ED) urgent care spaces of a regional medical center. Patients who triggered multi-parametric alerts [MPA] consisting of two or more red alarms (critical cardiac rhythm, heart rate [HR], blood pressure [BP], or oxygen saturation) within a 15-minute window were consented and matched into 1:1:1 study triads with patients who only triggered standard, single parameter (red) alarms [SPA] and those who remained in a non-alarm state [NA]. Study patients’ red alarms and triggering physiologic abnormalities (vital signs; cardiac rhythm and pulse oximetry photoplethysmography waveforms) were adjudicated and annotated for validity (chart review) and interpretability (clinician gestalt). Each subject was then followed for 3 months via in-network medical records. Subjects’ characteristics, alarms, clinical courses, and outcomes were descriptively and comparatively analyzed. The relationship between alarm-triggering waveform interpretability and alarm validity was analyzed with correlational statistics, along with the association between multi-parametric alerts and severity of patient state.Of 264 ED patients enrolled over 24 months, 183 matched into 61 triads matched for sex, age, Emergency Severity Index (ESI), and chief complaint; 27.1% of MPA, 10.9% of SPA, and 5.2% of NA subjects started care in ED critical care (X2(2)=6.47, p=0.04). MPA and SPA subjects triggered single parameter red alarms of all types, with a predominance of tachycardia and hypoxia alarms. Asystole and ventricular fibrillation/ventricular tachycardia (VFib/VTach) alarms were mostly adjudicated as non-valid, and hypoxia alarms were frequently of indeterminate validity, whereas HR and BP alarms were generally valid. Five (8.2%) MPA subjects and no SPA subjects (and no NA subjects) were transferred to ED critical care while monitored in the study area, X2(1)=9.41, p=0.002). Eleven (18.1%), five (8.2%), and two (3.3%) MPA, SPA, and NA subjects, respectively, were admitted from the ED to an intensive or intermediate care unit (p<0.001). Higher rates of symptomatic/unstable tachycardia (p=0.02) and post-ED escalation in care unit requirement (p=0.001) were recorded in MPA and SPA groups than in the NA group. Waveform interpretability scores in the 53 MPA subjects with complete data exhibited moderate positive correlation with the adjudicated validity of their red alarms, r(51)=0.528, p<0.001. The difference in composited 3-month tracer metrics between MPA and SPA groups did not attain statistical significance, although death (4.9%), Advanced Life Support (1.6%), and cardiac arrest (1.6%) were noted only in MPA patients.ED patient monitor datastreams were acquired and analyzed to explore novel biomedical engineering approaches to mitigate alarm fatigue. Safe experimentation in a live ED environment demonstrated the potential for near-real-time signals-level analysis of standard biomedical device outputs to differentiate true alarms from false alarms and to identify patients with a higher likelihood of poor outcomes.
The rapid spread of Monkeypox has underscored the importance of accurate and reliable diagnostic tools, particularly dermatological assessment. This research introduces a deep-learning method to accurately classify Monkeypox skin lesions, emphasizing transparency through model-agnostic explainability techniques. To achieve this, we incorporated Local Interpretable Model-agnostic Explanations (LIME) and SHapley Additive exPlanations (SHAP) into our analytical framework. We developed a Convolutional Neural Network (CNN) model to accurately distinguish Monkeypox skin lesions, utilizing a selected dataset of high-resolution lesion images. LIME was employed to generate detailed explanations for individual predictions, helping to pinpoint the exact features within the lesion images that most strongly influenced the model's decisions. Additionally, SHAP was utilized to assess the contribution of every feature to the model's overall predictions. It provides comprehensive information on the model's behaviour and ensures consistency during decision-making. The insights derived from these explainability methods were crucial in validating the model's reliability and interpreting misclassification instances, which informed subsequent model refinements. Quantitative results demonstrated that our model not only achieved a high accuracy of 92.19% but also concentrated on clinically significant areas of the lesions, as verified by LIME and SHAP visualizations. These findings suggest that integrating LIME and SHAP within deep learning frameworks can greatly enhance AI-powered diagnostic tools' trustworthiness and clinical relevance. Future research will aim to extend this approach to other dermatological conditions and investigate its potential integration into clinical decision-making processes.
Phishing emails are a significant threat to organizations, with over 90% of cyber attacks starting from a malicious email. Despite built-in security measures, relying solely on these defenses can leave organizations vulnerable to cybercriminals who exploit human nature and the lack of tight security. Phishing emails, designed to deceive recipients into disclosing personal and financial information, represent a significant cybersecurity challenge. This paper introduces a comprehensive dataset curated explicitly for detecting phishing emails, featuring a collection of authentic and phishing emails. The dataset includes a broad spectrum of phishing techniques, such as sophisticated social engineering tactics, impersonation of reputable entities, and using urgent or threatening language to manipulate recipients. Phishing emails were collected to cover various scenarios, including financial fraud, account verification, and malware dissemination attempts. Our analysis involves a range of classical machine learning models alongside exploratory analysis with LLMs. The performance of these models was rigorously evaluated to furnish a comparative analysis of their detection capabilities. The dataset, one of the largest of its kind, offers a significant resource for researchers and cybersecurity professionals aiming to advance phishing detection methods. The dataset used in this research is publicly available, enabling further exploration and replication of the findings by the research community [1].
Climate prediction is a complex and challenging task. Machine learning techniques, particularly Linear Regression combined with Convex Optimization, provide a promising alternative to traditional statistical models by improving predictive accuracy and reducing error. This research explores the application of these techniques to enhance the reliability of climate predictions. To ensure robustness and improve generalizability, the model was tested across multiple datasets, including Vostok Ice Core and ERA5, which capture historical and modern climate data. The results demonstrate a significant reduction in prediction error, highlighting the effectiveness of Convex Optimization in improving prediction accuracy and the model’s capacity to generalize across different climate datasets.