
This research introduces a novel approach for deterring woodpeckers to safeguard wooden structures in an environmentally-friendly way. Woodpecker drilling and drumming poses a significant threat to wooden structures, causing extensive physical and financial damage. These issues range from property damage to noise, underscoring the urgent need for an effective deterrent method. Our study proposes a laser-based deterrent mechanism, designed to be eco-friendly by minimizing effects on other wildlife and humans. This environmentally conscious approach offers a humane solution to woodpecker control, ensuring compliance with the Migratory Bird Treaty Act (MBTA). Additionally, our approach utilizes deep learning technology to chase away specific targets. Our system integrates YOLOv8 for accurate and low-power identification, creating an efficient operating solution with independent internet connectivity. This enhancement enables the system to adapt across various environments. Testing in a controlled environment has verified the system's effectiveness in deterring woodpeckers in an eco-friendly way. This approach ensures minimal impact on other wildlife and humans, adhering to legal and environmental constraints.
We present a multi-agent system simulation designed for efficient coordination and collaboration among multiple robots, particularly suited for search operations. This simulation reflects unstructured and complex outdoor scenarios where significant obstruction and occluded terrain surfaces cause difficulties in search. The software integrates well with reinforcement learning (RL) and a centralized Multi-Agent Transformer (MAT) to enable autonomous robots to collect, process, and integrate data. Search, coverage, and complex mobility planning can be tested in this simulation with dynamic and unstructured environments. The project code and videos can be found at https://github.com/DIRECTLab/Coordinating-MAT-Env
Infrastructure inspection, especially for wind turbines, skyscrapers, bridges, and dams, presents significant challenges due to the complexity of modern architecture and the hazardous environments involved. Current inspection methods require substantial human resources and equipment like scaffolding, cranes, and aerial lifts, which are costly and pose safety risks. Existing robotic solutions, such as drones and wheeled or tracked robots, are often limited in their ability to navigate vertical surfaces or carry significant loads. This research introduces VertiPaw, the first quadrupedal robot designed specifically for vertical climbing on various surfaces using a vacuum suction system with heavy loads. VertiPaw's adaptive spring limbs provide versatile movement on both flat and uneven surfaces, while its lightweight design ensures lower power consumption compared to similar robots. Equipped with Light Detection and Ranging (LIDAR) for precise positioning and surface detection, VertiPaw is designed to perform infrastructure inspections in difficult-to-reach areas. The system demonstrates efficient navigation across multiple surface types, including metal, concrete, and uneven walls, while maintaining stability and carrying substantial loads. VertiPaw offers enhanced adaptability, power efficiency, and safety. Its ability to perform complex tasks such as cargo movement and real-time surveillance makes it a valuable system for further development.
Real-time vision applications such as object detection for autonomous navigation have recently witnessed the emergence of neuromorphic or event cameras, thanks to their high dynamic range, high temporal resolution and low latency. In this work, our objective is to leverage the distinctive properties of asynchronous events and static texture information of conventional frames. Towards this goal, asynchronous events are first transformed into a 2D spatial grid representation, which is carefully selected to harness the high temporal resolution of event streams and align with conventional image-based vision. Via a joint detection framework, detections from both RGB and event modalities are fused by probabilistically combining scores and bounding boxes. The superiority of the proposed method is demonstrated over concurrent Event-RGB fusion methods on DSEC-MOD and PKU-DDD17 datasets.
Talk WithMachines aims to enhance human-robot interaction in safety-critical industrial systems by integrating large/vision language models with robot control and perception. This allows robots to understand natural language commands and perceive their environment. Translating robots' internal states into human-readable text allows operators to gain clearer insights for safer operations. The paper outlines four workflows: low-level control, language-based feedback, visual input, and robot structure-informed task planning, which are presented in a set of experiments. The proposed approach outperforms the prior method in grasping (100% success vs. 90%) and obstacle avoidance (50% success vs. 30%). Supplementary materials are available on the project website: https://talk-machines.github.io.
Recent advancements in Artificial Intelligence (AI) methodologies, particularly with the rise of Generative AI (GenAI), have paved the way for a new chapter in the robotics domain. In this short paper, we discuss the benefits of integrating GenAI into the robotics domain. Specifically, we analyze the application of GenAI-powered robotics technologies in industrial inspection, defining the specific needs and proposing an approach that best leverages the characteristics of GenAI models to address these challenges in robotic industrial inspection scenarios.
Stampede accidents frequently occur in situations where large crowds gather. Although education on how to protect oneself and prevent accidents in dense crowds has been provided, these measures mainly focus on post-accident responses. In contrast, this study proposes a proactive approach to prevent stampede accidents by utilizing thermal cameras to detect the number of people in a space in real-time and calculate the risk of a stampede. The system collects object detection and density estimation results using thermal cameras, considering population density in a specific area, and transmits the estimated results to other devices via serial communication. Thermal imaging technology excels at detecting people with high accuracy even in challenging daytime conditions or low-light environments at night. The data collected from the thermal cameras is continuously updated through machine learning and pattern analysis to assess stampede risk, and the results are provided in real-time via a web interface. This allows safety personnel and managers to effectively monitor high-density areas and take immediate action if necessary. Additionally, the system's web interface provides users with real-time information related to stampede risk. The validity of the proposed system has been demonstrated through field applicability and performance evaluation in real-world environments, contributing to enhanced public safety by preventing stampede accidents in advance.
This research addresses the security challenges posed by Generative Adversarial Networks (GANs) in biometric authentication, focusing particularly on the use of Hyperspectral Imaging (HSI) to counteract the threat of deep fake biometrics created by GANs. We sought solutions to the issues identified in a previous study, such as the inability to handle images of individuals not previously trained and the high time consumption required for identification. In this paper, we first reduced the dimensionality of hyperspectral face data used for training, and then conducted a binary classification of “target subject (class)” and “nontarget subjects (classes).” A discriminator, based on the Least Squares Generative Adversarial Network (LSGAN) model, was developed using HSI for each class for binary classification. The proposed method significantly reduced identification time and achieved very high accuracy in identification.
The integration of Human-Robot Collaboration (HRC) into industrial production aims to enhance efficiency by combining human intuition with robotic precision. Traditional lab environments often fail to capture the complexity of real-world settings, hindering effective HRC research. This study addresses this gap by using a 3D cave setup with stereoscopic 360-degree video and dynamic Unity 3D content to create a more immersive experimental environment. Twelve participants tested an industrial robot in three conditions: no cave, non-responsive 3D video, and interactive 3D content. Results showed that the immersive setup improved role assumption, interaction, and contextual awareness, bridging the gap between lab research and real-world applications.
Achieving a stable displacement of an object to its target pose is a fundamental requirement in numerous robotic manipulation tasks. In this work, we investigate the use of planar pushing to re-orient an object with unknown physical parameters. Primarily, this study serves as a supplement to our previously introduced Zero Moment Two Edge Pushing (ZMTEP) technique designed to achieve pure object translation. Specifically, precise object re-orientation is accomplished by employing a variable stroke parallel-jaw gripper pusher to make contact with two edges of the object and execute circular motion. We initially give an assertion for enabling an unknown object to smoothly track a specified circular trajectory while remaining in sticking contact with the pusher. Subsequently, we carry out a comprehensive set of experiments to validate the suggested claim by examining the possible two-edge-contact (TEC) configurations. Lastly, we assess the practicality of the estimated frictional forces to find the achievable TEC configurations. The experimental outcomes provide empirical evidence that confirms the validity of friction estimation for TEC pushing.
Although exploration methods for homogeneous multi-robot systems have improved, they generally lack consideration for the impact of robot capabilities on exploration in complex environments. To address this, we propose a task-allocation method based on Bayesian Networks for decentralized heterogeneous systems. Tasks are randomly distributed in unexplored areas, guiding robots to complete the exploration. Our method is demonstrated in real-world ground environments.
As Unmanned Aerial Vehicles (UAVs) become more accessible to the public, they become a common tool for malicious purposes. As a result, there is an increasing demand for Counter Unmanned Aerial Systems (CUAS) that can detect UAVs. Existing CUAS solutions often rely on high-priced radar systems or advanced technologies, primarily designed for military purposes. In this paper, a low-cost, effective, non-military CUAS that uses inexpensive smartphones' microphones and camera, along with machine learning models, is proposed to detect and track a malicious UAV (MUAV) in real-time. Our proposed CUAS is designed to be affordable and accessible to the general public, operating automatically to detect and track MUAVs in real-time.
The rapid development of unmanned aerial vehicles (UAVs) has intensified the need for advanced classification techniques. This paper presents a novel approach that leverages audio data transformed into visual representations through Mel-Frequency Cepstral Coefficients (MFCCs) for drone classification. Our dataset consists of 28 drone types, each with 100 five-second audio recordings, from which 30 MFCCs are extracted per file. We investigate the effectiveness of this dataset by applying various vision models to the MFCC visualizations. Our results reveal that EfficientNet achieved the highest accuracy at 96.31%, followed by ResNet50 at 94.22%, and Vision Transformer at 73.69%. These findings highlight the potential of using audio-derived visual features for robust drone classification and demonstrate the varying performance of different vision models. This study provides a comprehensive examination of the methodology, experimental setup, and results, offering valuable insights into future research directions for enhancing classification accuracy with transformed audio data.
In Human-Robot Interaction, the robot's ability to recognize human frustration signals and respond appropriately is crucial. Large Language Models (LLMs) are increasingly used within conversational agents and social robots. In this preliminary study we test four LLMs by simulating three different scenarios: health care, education, and public administration. The goal is twofold: first, we develop a qualitative and quantitative conversation analysis methodology, and second, we test the ability of these LLMs to recognize and respond appropriately to frustration in different contexts.
The use of robotics and generative AI has the potential to increase automation of industrial inspection processes, leading to improved asset management and significant cost savings. This paper describes how vision language models can be used to process image data acquired by robots to perform inspection and streamline the maintenance response. The application of this technology is considered with respect to deploying robotics for visual inspection of railway assets. Whilst there are some industry specific constraints and benefits, the general concepts presented should be readily applicable to other asset intensive industries.
The ability of robots to imitate human learning strategies-rapidly adapting to new tasks without large datasets-has garnered significant attention in meta-learning. Meta-reinforcement learning seeks to enhance robotic agent flexibility across diverse tasks and contexts, offering promise where single-task learning often fails. Despite advancements like multi-task diffusion models and task-weighted optimization mechanisms, effectively training tasks with varying complexities simultaneously remains a major challenge. This paper introduces a novel meta-reinforcement learning method that addresses this issue by clustering the training tasks of robotic arms based on semantic and trajectory similarities, while leveraging adaptive learning rates and task-specific weights proposed by the multitask optimization techniques. Our approach, TEAM, emphasizes performance-driven semantic clustering, optimizing based on robotic task similarity, complexity, and convergence objectives. We also integrate fast adaptive and multi-task optimization of the diffusion model to enhance computational efficiency and adaptability. More specifically, we introduce a cluster-specific optimization technique, using specialized parameters for each group to allow more refined task handling. The experimental validation demonstrates the effectiveness of this scalable method in improving performance, adaptability, and efficiency in real-world, heterogeneous robotic tasks, further advancing robotic computing in meta-reinforcement learning.