Pixel-level annotation remains a major bottleneck for semantic segmentation, motivating methods that synthesize image–label pairs directly from generative models. Prior synthetic dataset generators typically obtain pseudo-labels from cross-attention maps or learned decoders over generative features; however, recent text-to-image (T2I) models increasingly use multimodal diffusion transformers (MM-DiTs), where concept localization is no longer exposed through a single cross-attention pathway but instead distributed across many layers and attention heads. Existing MM-DiT localization methods address this by aggregating saliency across heads, but we observe that this averaging can dilute clean target localizers due to attention head heterogeneity. We introduce HEADHUNTER, a training-free segmentation framework that uses aggregate concept saliency as a self-guided proxy to select the single attention head that best localizes a queried textual concept, yielding cleaner segmentation masks. We then use HEADHUNTER to turn target classes into training data automatically: a large language model (LLM) diversifies prompts, an MM-DiT generates images, HEADHUNTER produces pseudo-labels, and a vision language model (VLM) verifies each image–mask pair before acceptance. HEADHUNTER achieves strong zero-shot segmentation performance (81.2 mIoU on PASCAL VOC2012 and 73.2 mIoU on ImageNet-Segmentation), outperforming head aggregation and other interpretability methods. Using only our generated image–label pairs, we train segmentation models which reach 64.9 mIoU on VOC2012 validation, matching or outperforming comparable synthetic dataset generators and showing that the proposed pipeline produces effective dense labels without human intervention.
Dynamic object detection is fundamental to advancing vision-based navigation systems, particularly in environments where the camera itself is in motion. Despite progress in detection algorithms, existing approaches often struggle with challenges such as egomotion, short-term occlusions, temporal discontinuities, and computational cost. This paper presents YOLO-MOTF, a novel knowledge-based model that integrates spatial features with motion cues, especially for operation under moving camera conditions. The framework incorporates a hybrid motion compensation strategy to suppress camera-induced distortions and an occlusion handling buffer to preserve object trajectories through discontinuities. Additionally, a motion attention gating mechanism selectively reinforces moving object predictions by intersecting fused motion masks with semantic outputs. The proposed system achieves an F1 score of 88.6% and a 93% reduction in flow processing compared to dense flow methods, underscoring its robustness and efficiency in dynamic environments. Beyond theoretical contributions, the model demonstrates direct applicability in real-world knowledge-based decision systems, including healthcare applications such as assistive wheelchair navigation, as well as assistive robotics, autonomous navigation, and surveillance.
Object detection plays a pivotal role in advancing computer vision systems by enabling machines to perceive and interact intelligently with their environments. Despite significant advancements, comprehensive exploration of its evolution and applications in navigation remains underrepresented. This review paper examines the evolution of object detection technologies, from early methodologies to contemporary advancements, and their critical role in navigation tasks. The emphasis was on the significance of contextual learning in enhancing object detection performance by leveraging spatial and temporal information. Furthermore, the limitations of conventional approaches that rely heavily on hand-engineered features are examined. It is then demonstrated that contextual learning facilitates automated feature extraction, resulting in improved accuracy exceeding a 50% increase and adaptability in diverse applications. The review concludes by outlining future trends and opportunities for further advancements in object detection and, underscoring its transformative impact on autonomous navigation and beyond. In summary, this review contributes to a comprehensive understanding of object detection technologies by offering insights into their evolution, highlighting their applications in navigation, and providing guidance for future research in context-aware systems.
Navigating through doorways remains a daily challenge for wheelchair users, often leading to frustration, collisions, or dependence on assistance. These challenges highlight a pressing need for intelligent doorway detection algorithm for assistive wheelchairs that go beyond traditional object detection. This study presents the algorithmic development of a lightweight, vision-based doorway detection and alignment module with contextual awareness. It integrates channel and spatial attention, semantic feature fusion, unsupervised depth estimation, and doorway alignment that offers real-time navigational guidance to the wheelchairs control system. The model achieved a mean average precision of 95.8% and a F1 score of 93%, while maintaining low computational demands suitable for future deployment on embedded systems. By eliminating the need for depth sensors and enabling contextual awareness, this study offers a robust solution to improve indoor mobility and deliver actionable feedback to support safe and independent doorway traversal for wheelchair users.
Slope failures in open-pit mining pose significant operational and safety issues, underscoring the importance of implementing effective stability monitoring frameworks for early hazard detection to allow for timely intervention and risk mitigation. This systematic review presents a comprehensive synthesis of existing and emerging methods and technologies used for slope stability monitoring in open-pit mining, including both remote sensing and in situ methods, as well as advanced technologies, such as Artificial Intelligence (AI), the Internet of Things (IoT), and Wireless Sensor Networks (WSNs). Using the Preferred Reporting Items for Systematic Reviews and Meta-Analysis (PRISMA) 2020 guidelines, a total of 49 studies were selected from a collection of four engineering databases, and a comparative analysis was conducted to determine the underlying differences between the various methods for open-pit slope stability monitoring in terms of their performance across key attributes, such as monitoring accuracy, spatial and temporal coverage, operational complexity, and economic viability. Their juxtaposition highlighted the notion that no universally optimal slope stability monitoring system exists, due to a series of compromises that arise as a result of inherent technological limitations and site-specific constraints. Notably, remote sensing methods offer large-scale, non-intrusive monitoring, but are often limited by environmental factors and data acquisition infrequency, whereas in situ methods provide high precision, but suffer from limited spatial coverage and scalability. This review further highlights the capacity of emerging methods and technologies to address these limitations, providing suggestions for future research directions involving the integration of multiple sensing technologies for the enhancement of monitoring capabilities. This study provides a consolidated knowledge base on open-pit slope stability monitoring methods, technologies, and techniques, to guide the development of integrated, cost-effective, and scalable slope monitoring solutions that enhance mine safety and efficiency.
Advancements in Unmanned Aerial Vehicle (UAV) technologies have increased exponentially in recent years, with UAV swarm being a key area of interest. UAV swarm overcomes the energy reserve, payload, and single-objective limitations of single UAVs, enabling broader mission scopes. Despite these advantages, UAV swarm has yet to see widespread application within global industry. A leading factor hindering swarm application within industry is the divide that currently exists between the functional capacity of modern UAV swarm systems and the functionality required by legislation. This paper investigates this divide through an overview of global legislative practice, contextualized via a case study of Australia’s UAV regulatory environment. The overview highlighted legislative objectives that coincided with open challenges in the UAV swarm literature. These objectives were then formulated into analysis criteria that assessed whether systems presented sufficient functionality to address legislative concern. A systematic review methodology was used to apply analysis criteria to multi-objective UAV swarm mission planning systems. Analysis focused on multi-objective mission planning systems due to their role in defining the functional capacity of UAV swarms within complex real-world operational environments. This, alongside the popularity of these systems within the modern literature, makes them ideal candidates for defining new enabling technologies that could address the identified areas of weakness. The results of this review highlighted several legislative considerations that remain under-addressed by existing technologies. These findings guided the proposal of enabling technologies to bridge the divide between functional capacity and legislative concern.
In this paper, we introduce the Secured Real-Time Machine Communication Protocol (SRMCP), a novel industrial communication protocol designed to address the increasing demand for security and performance in Industry 4.0 environments. SRMCP integrates post-quantum cryptographic techniques, including the Kyber Key Encapsulation Mechanism (Kyber-KEM) and AES-GCM encryption, to ensure robust protection against both current and future cryptographic threats. We also present an innovative “Port Hopping” mechanism inspired by frequency hopping, enhancing security by distributing communication across multiple channels. Comparative performance analysis was conducted with widely-used protocols such as ModBus and the OPC UA, focusing on key metrics such as connection, reading, and writing times across local and remote networks. Results demonstrate that SRMCP outperforms ModBus in reading and writing operations while offering enhanced security, although it has a higher connection time due to its dual-layer encryption. The OPC UA, while secure, lags significantly in performance, making it less suitable for real-time applications. The findings suggest that SRMCP is a viable solution for secure and efficient machine communication in modern industrial settings, particularly where quantum-safe security is a concern.
The growing presence of autonomous vehicles (AVs) on our roads has been applauded by many people, including those in the transportation, military, and logistics industries. However, an equally growing concern about AVs is that they raise new cybersecurity challenges, including the need for a robust digital forensic investigation framework to investigate incidents or criminal activities involving AVs. Of more concern is that traditional digital forensic methods have been found wanting when investigating incidents involving AVs. This is primarily due to the complex interplay of sensors, software, and human interaction involved in AV operation. In this research, therefore, the authors review existing digital forensic investigation frameworks for AVs and then analyse their strengths, weaknesses, and suitability for addressing various AV forensic challenges. Additionally, we analyse the potential legal and ethical implications of digital forensic investigations in AV contexts. The authors aim to identify key AV forensic challenges and future directions through this research. As a new contribution, the authors propose high-level ontological digital forensic investigation framework components as a starting point towards a comprehensive framework and the basis for further exploration and future research. This includes the future implementation of the proposed components into a fully-fledged ontological framework that supports efficient and reliable digital forensic investigation processes for autonomous vehicles.
Models of seven discrete facial expressions are built on macro-level facial muscle variations for separating distinct affective states. We propose a step-wise Hierarchical Separation and Classification Network (HSCN) that discovers dynamic and continuous macro- and micro-level variations in facial expressions. The HSCN first invokes an unsupervised cosine similarity-based separation method on continuous facial expression data and extracts twenty-one dynamic expression classes from the seven common discrete affective states. Separation between the clusters is then optimised for discovering the macro-level changes in facial muscle activations followed by splitting the upper and lower facial regions for realising and modelling changes pertaining to upper and lower facial muscle activations. A linear discriminant space is developed for clustering the upper and lower facial images on the basis of similar muscular activation patterns. Actual dynamic data and linear discriminant features are mapped for developing a rule-based expert system that would facilitate classification of twenty-one upper and twenty-one lower facial micro-expressions. Using the random forest algorithm, classification accuracies of 76.11\% were observed for dynamic macro-level facial expression classification. A support vector machine provided 73.63\% and 87.68\% accuracies respectively while classifying upper and lower facial micro-expressions. This work provides a novel framework for the dynamic assessment of affective states. Reported methods and results also provide new insight into the dynamic analysis of facial expressions of affective states.
Objective Route planning is key support, which can be provided by navigation tools for the blind and visually impaired (BVI) persons. No comprehensive analysis has been reported in the literature on this topic. The main objective of this study is to examine route planning approaches used by indoor navigation tools for BVI persons with respect to the route determination criteria, how user context is integrated, algorithms adopted, and the representation of the environment for route planning. Methods A systematic review was conducted covering the period 1982-July 2018 and thirteen databases. Results From the 1709 articles that resulted from the initial search, 131 were selected for the study. Route length was the sole factor used to determine the best route in the majority of the studies. Routes with less obstacles, less turns, more landmarks and close to walls were selected by the other studies. User variations were considered in few studies. Unique needs of the BVI persons were addressed by few novel algorithms which had integrated multiple parameters and BVI-friendly route modelling. Conclusion Differences between the navigation capabilities of the sighted and the BVI community were not a major concern when deciding the optimum routes in the majority of BVI indoor navigation tools. How to trade-off between factors affecting optimum routes, and how to model suitable routes in complex buildings needs to be studied further, looking from the user perspective.
Gesture recognition is a mechanism by which a system recognizes an expressive and purposeful action made by a user’s body. Hand-gesture recognition (HGR) is a staple piece of gesture-recognition literature and has been keenly researched over the past 40 years. Over this time, HGR solutions have varied in medium, method, and application. Modern developments in the areas of machine perception have seen the rise of single-camera, skeletal model, hand-gesture identification algorithms, such as media pipe hands (MPH). This paper evaluates the applicability of these modern HGR algorithms within the context of alternative control. Specifically, this is achieved through the development of an HGR-based alternative-control system capable of controlling of a quad-rotor drone. The technical importance of this paper stems from the results produced during the novel and clinically sound evaluation of MPH, alongside the investigatory framework used to develop the final HGR algorithm. The evaluation of MPH highlighted the Z-axis instability of its modelling system which reduced the landmark accuracy of its output from 86.7% to 41.5%. The selection of an appropriate classifier complimented the computationally lightweight nature of MPH whilst compensating for its instability, achieving a classification accuracy of 96.25% for eight single-hand static gestures. The success of the developed HGR algorithm ensured that the proposed alternative-control system could facilitate intuitive, computationally inexpensive, and repeatable drone control without requiring specialised equipment.
Personal protective equipment (PPE) is an essential key factor in standardizing safety within the workplace. Harsh working environments with long working hours can cause stress on the human body that may lead to musculoskeletal disorder (MSD). MSD refers to injuries that impact the muscles, nerves, joints, and many other human body areas. Most work-related MSD results from hazardous manual tasks involving repetitive, sustained force, or repetitive movements in awkward postures. This paper presents collaborative research from the School of Electrical Engineering and School of Allied Health at Curtin University. The main objective was to develop a framework for posture correction exercises for workers in hostile environments, utilizing inertial measurement units (IMU). The developed system uses IMUs to record the head, back, and pelvis movements of a healthy participant without MSD and determine the range of motion of each joint. A simulation was developed to analyze the participant’s posture to determine whether the posture present would pose an increased risk of MSD with limits to a range of movement set based on the literature. When compared to measurements made by a goniometer, the body movement recorded 94% accuracy and the wrist movement recorded 96% accuracy.
Internationally, the recent pandemic caused severe social changes forcing people to adopt new practices in their daily lives. One of these changes requires people to wear masks in public spaces to mitigate the spread of viral diseases. Affective state assessment (ASA) systems that rely on facial expression analysis become impaired and less effective due to the presence of visual occlusions caused by wearing masks. Therefore, ASA systems need to be future-proofed and equipped with adaptive technologies to be able to analyze and assess occluded facial expressions, particularly in the presence of masks. This paper presents an adaptive approach for classifying occluded facial expressions when human faces are partially covered with masks. We deployed an unsupervised, cosine similarity-based clustering approach exploiting the continuous nature of the extended Cohn-Kanade (CK+) dataset. The cosine similarity-based clustering resulted in twenty-one micro-expression clusters that describe minor variations of human facial expressions. Linear discriminant analysis was used to project all clusters onto lower-dimensional discriminant feature spaces, allowing for binary occlusion classification and the dynamic assessment of affective states. During the validation stage, we observed 100% accuracy when classifying faces with features extracted from the lower part of the occluded faces (occlusion detection). We observed 76.11% facial expression classification accuracy when features were gathered from the uncovered full-faces and 73.63% classification accuracy when classifying upper-facial expressions - applied when the lower part of the face is occluded. The presented system promises an improvement to visual inspection systems through an adaptive occlusion detection and facial expression classification framework.
The majority of human gait modeling is based on hip, foot or thigh acceleration. The regeneration accuracy of these modeling approaches is not very high. This paper presents a harmonic approach to modeling human gait during level walking based on gyroscopic signals for a single thigh-mounted Inertial Measurement Unit (IMU) and the flexion–extension derived from a single thigh-mounted IMU. The thigh angle can be modeled with five significant harmonics, with a regeneration accuracy of over 0.999 correlation and less than 0.5° RMSE per stride cycle. Comparable regeneration accuracies can be achieved with nine significant harmonics for the gyro signal. The fundamental frequency of the harmonic model can be estimated using the stride time, with an error level of 0.0479% (±0.0029%). Six commonly observed stride patterns, and harmonic models of thigh angle and gyro signal for those stride patterns, are presented in this paper. These harmonic models can be used to predict or classify the strides of walking trials, and the results are presented herein. Harmonic models may also be used for activity recognition. It has shown that human gait in level walking can be modeled with a harmonic model of thigh angle or gyro signal, using a single thigh-mounted IMU, to higher accuracies than existing techniques.
Background Cerebral palsy (CP) is a physical disability that affects movement and posture. Approximately 17 million people worldwide and 34,000 people in Australia are living with CP. In clinical and kinematic research, goniometers and inclinometers are the most commonly used clinical tools to measure joint angles and positions in children with CP. Objective This paper presents collaborative research between the School of Electrical Engineering, Computing and Mathematical Sciences at Curtin University and a team of clinicians in a multicenter randomized controlled trial involving children with CP. This study aims to develop a digital solution for mass data collection using inertial measurement units (IMUs) and the application of machine learning (ML) to classify the movement features associated with CP to determine the effectiveness of therapy. The results were calculated without the need to measure Euler, quaternion, and joint measurement calculation, reducing the time required to classify the data. Methods Custom IMUs were developed to record the usual wrist movements of participants in 2 age groups. The first age group consisted of participants approaching 3 years of age, and the second age group consisted of participants approaching 15 years of age. Both groups consisted of participants with and without CP. The IMU data were used to calculate the joint angle of the wrist movement and determine the range of motion. A total of 9 different ML algorithms were used to classify the movement features associated with CP. This classification can also confirm if the current treatment (in this case, the use of wrist extension) is effective. Results Upon completion of the project, the wrist joint angle was successfully calculated and validated against Vicon motion capture. In addition, the CP movement was classified as a feature using ML on raw IMU data. The Random Forrest algorithm achieved the highest accuracy of 87.75% for the age range approaching 15 years, and C4.5 decision tree achieved the highest accuracy of 89.39% for the age range approaching 3 years. Conclusions Anecdotal feedback from Minimising Impairment Trial researchers was positive about the potential for IMUs to contribute accurate data about active range of motion, especially in children, for whom goniometric methods are challenging. There may also be potential to use IMUs for continued monitoring of hand movements throughout the day. Trial Registration Australian New Zealand Clinical Trials Registry (ANZCTR) ACTRN12614001276640, https://www.anzctr.org.au/Trial/Registration/TrialReview.aspx?id=367398; ANZCTR ACTRN12614001275651, https://www.anzctr.org.au/Trial/Registration/TrialReview.aspx?id=367422
Cerebral palsy (CP) is a common reason for human motor ability limitations caused before birth, through infancy or early childhood. Poor head control is one of the most important problems in children with level IV CP and level V CP, which can affect many aspects of children’s lives. The current visual assessment method for measuring head control ability and cervical range of motion (CROM) lacks accuracy and reliability. In this paper, a HeadUp system that is based on a low-cost, 9-axis, inertial measurement unit (IMU) is proposed to capture and evaluate the head control ability for children with CP. The proposed system wirelessly measures CROM in frontal, sagittal, and transverse planes during ordinary life activities. The system is designed to provide real-time, bidirectional communication with an Euler-based, sensor fusion algorithm (SFA) to estimate the head orientation and its control ability tracking. The experimental results for the proposed SFA show high accuracy in noise reduction with faster system response. The system is clinically tested on five typically developing children and five children with CP (age range: 2–5 years). The proposed HeadUp system can be implemented as a head control trainer in an entertaining way to motivate the child with CP to keep their head up.
Gait modelling is essential for many applications including animation, activity recognition, medical diagnosis, and robotics. Many researchers have worked on mathematically express the movement of human bodies. At the current stage, the reconstructed waveforms from the mathematical expressions either represent smoothened waveforms, noisy, or require a high number of computations. In this study, the thigh and shank angle waveforms are time and amplitude scaled before performing a discrete Fourier transform (DFT). By doing so, the correlation coefficient between the original and reconstructed waveforms can be improved without increasing the number of harmonics. The shank's angular velocity is also recalculated from the reconstructed shank's angle waveform for gait phase detection, and shows accurate results in heel and toe strikes estimation when compared to the original shank's angular velocity. Additionally, the harmonic components of the waveforms are used for gait recognition. Experimental results show that it is useful to time and amplitude-scale the angle waveforms to ‘enlarge’ the distinctive regions of the angle waveforms for better classification accuracy.
In this research, the authors investigate the feasibility of selecting three-dimensional thigh and shank angles as the features of machine learning methods. Four common machine learning techniques, i.e. random forest, k-nearest neighbour, support vector machine and perceptron, were compared in terms of accuracy and memory usage so that a real-time standalone gait diagnosis device can be constructed using low-end inertial measurement units (IMUs). With proper re-sampling and normalisation, they discovered that the support vector machine and perceptron resulted in the top two highest accuracies (96-99%) among the four machine learning methods. The memory requirement of the perceptron is the lowest among the machine learning methods. Therefore, perceptron was selected as the classification algorithm for the standalone gait diagnosis device. The trained perceptron was transferred to the thigh and shank's IMUs to process the data locally in real-time. The constructed standalone gait diagnosis device lit up green or red light emitting diodes when normal or abnormal gaits were detected, respectively. This standalone device was further tested in real-life and achieved a mean classification accuracy of 96.50%.
In the above paper [1], Table I is corrected as follows.
Recent statistics of the World Health Organization (WHO) indicate that over 253 million of the world’s population to be visually impaired. Most of these individuals use the white cane as an assistive tool or are often accompanied by caretakers or voluntary helpers as indoor navigation is particularly challenging for them. This chapter describes a substitute vision system designed to assist vision impaired individuals through the use of visible light communication and geomagnetism. Furthermore, the use of database optimization increases the speed and efficiency of data retrieval thus reducing system response time. Though navigation systems that support the visual impaired are readily available, there have been no systems that use both visible light communication and geomagnetism capable of providing accurate and secure indoor navigation assistance, which in turn would increase the overall satisfaction of the system users.
Tele Tan合作论文数School of Civil and Mechanical Engineering, Faculty of Science and Engineering, Curtin University3