Due to the highly uncertainty of pedestrian movement on the road, understanding pedestrian behavior remains a complex and challenging task in the field of autonomous driving. While numerous studies have been conducted on pedestrian action recognition for intention prediction, we focus on pedestrian gesture detection, an aspect that has received little attention. Pedestrians often use certain arm gestures to express their intentions to oncoming vehicles. Detection of pedestrian gestures assists autonomous vehicles in understanding pedestrian intentions like human drivers. We propose a skeleton-based approach for pedestrian arm gesture detection and recognition. A pose estimation algorithm is applied to extract skeleton points of pedestrians. The angle and relative position between the pedestrian's arm and body are extracted for online arm gesture detection. Then the angle and moving pose descriptor is adopted for gesture recognition with support vector machines as the classifier. Experimental results on multiple datasets show that our proposed method outperforms other similar work.
Understanding pedestrian behaviors is significant for safe interaction between autonomous vehicles and pedestrians. Though numerous achievements have been reached for autonomous driving, this is still an open issue due to the highly uncertainty and dynamicity of pedestrians. Early studies treated pedestrians as moving rigid objects, predicting their future positions by trajectory prediction. More recent studies attempted to recognize pedestrian intentions by their body pose and actions. However, the intention of pedestrians is simply recognized as a binary result, cross or not-cross. This is not sufficient to describe the dynamic communication process between pedestrians and vehicles. We argue that pedestrian behaviors should be further explored and interpreted instead of concluding whether or not to cross, especially when pedestrians use hand gesture to communicate with vehicles. Hence, we established a taxonomy of pedestrian interactive behaviors. A new large-scale dataset, involving nine types of behaviors and 3600 video clips, is published. We adopted pose estimation to obtain 2D key points on pedestrian skeleton. A covariance descriptor is applied to depict high-level spatio-temporal features based on body pose, which is independent of sequence length and robust to missing key points. Then pedestrian behaviors are interpreted using Random Forest. Experiments both on the proposed dataset and publicly available JAAD dataset proved that the proposed method outperformed the state-of-the-art by approximately 5%.
Linear friction welding (LFW) was conducted on TC11 (Ti-6.5Al-3.5Mo-1.5Zr-0.3Si) and TC17 (Ti-5Al-2Sn-2Zr-4Mo-4Cr) titanium alloys shaped in the rectangular section of 75 mm × 20 mm, in which the welding parameters were set as oscillation frequency 20–60 Hz with amplitude 2–3 mm, friction pressure 58.8–65.3 MPa with friction time 3–5 s, and upset pressure 49–78.4 MPa with upset time 30 s. The microstructure characteristics and fatigue limit were measured and analyzed. After LFW, a microstructure change occurs with martensite in the weld zone and elongated grains in the thermo-mechanically affected zone. The effect of amplitude and frequency on fatigue properties is larger than that of friction time and forging pressure. With the amplitude and frequency of oscillation increasing, the fatigue limit increases firstly and then decreases. Accordingly, the optimal parameters were gained. The change tendency of α phase at failure location with the welding parameters could well describe the fatigue limit. Meanwhile, the fracture surface of joints presents a composite fracture of the quasi-cleavage, fatigue striations, and dimples. The spacing of fatigue striations firstly becomes smaller and then larger with the amplitude increasing which agrees with the change of fatigue limit.
Feature representation is of vital importance for human action recognition. In recent few years, the application of deep learning in action recognition has become popular. However, for action recognition in videos, the advantage of single convolution feature over traditional methods is not so evident. In this paper, a novel feature representation that combines spatial and temporal feature with global motion information is proposed. Specifically, spatial and temporal feature from RGB images is extracted by convolutional neural network (CNN) and long short-term memory (LSTM) network. On the other hand, global motion information extracted from motion difference images using another separate CNN. Hereby, the motion difference images are binary video frames processed by exclusive or (XOR). Finally, support vector machine (SVM) is adopted as classifier. Experimental results on YouTube Action and UCF-50 show the superiority of the proposed method.
This study aims to evaluate the high-temperature tensile and fatigue properties of FGH96 inertia friction welded (IFW) joints obtained using different welding parameters. The joint presented a significant microstructure change across the faying interface, characterized by the very small uniform equiaxed grains of the weld nugget zone (WNZ), obvious grain deformation, and growth of the thermo-mechanically affected zone (TMAZ). The elevated temperature tensile and high-cycle fatigue tests were performed at 700 °C. The results exhibit that the effect of rotational speed on tensile properties is less than that of friction pressure. The change tendency of tensile properties with welding parameters is in agreement with that of the width of WNZ. The tensile failure occurred in the WNZ, which is related to the complete-ordered Ni3Al γ′-phase dissolution. The fatigue limit decreases slightly along with the rotational speed. As the friction pressure increases, the fatigue limit increases firstly, and then decreases. The fatigue failure of the joint is located at the border between WNZ and TMAZ. This is related to the microhardness difference (dH) between TMAZ and WNZ, which reflects the stress concentration factor (Kt). In the end, the fracture mechanism was observed and analyzed.
Motion representation plays a vital role in human action recognition. In recent few years, the application of deep learning in action recognition has become popular. However, there are great challenges in extracting accurate motion features. In this study, a novel feature representation that combines multi-scale spatial-temporal feature is proposed. This descriptor contains spatial-temporal information for three mode, which are extracted from three input channels of RGB images, RGB difference images and binary XOR images. Specifically, a network that consist of convolutional neural network (CNN) and long short-term memory (LSTM) extract spatial-temporal feature from RGB images and RGB difference images respectively. On the other hand, global motion information is extracted from binary XOR images using another separate CNN network. Then, we combine this features from the three channels as a new video feature representation. Finally, an extreme learning machine (ELM) is adopted as classifier. Experimental results on UCF-50 dataset show the superiority of the proposed method.
Inertia friction welding (IFW) was conducted on FGH96 superalloy shaped in rings. After welding, the joints experienced solution treatment and aging-treated. There are 4 kinds of solution treatments including no solution treatment and solution treatment at 1050 C, 1080 C and 1120 C, respectively. The grain size, ?? precipitate distribution, tensile and fatigue properties at 700 C were measured and analyzed at different solution treatments. The results shows that with the increasing of the solution treatment temperature (STT), the grain size went up slowly below 1080 C and got a sudden growth from 1080 C to 1120 C. At the same time, the secondary ?? precipitates (mainly containing Ti, Al and Ni with ordered L1(2) structure) became fine and the shape gradually transformed from sphere into elliptical and irregular, which was a mixture of cuboids with concave faces and octets. When the solution treatment changes from no solution to solution treatment at 1120 C, the tensile strength gradually increases from 1243 MPa to 1337 MPa and then remains constant from 1080 C to 1120 C. However, the yield strength increases from 917 MPa to 990 MPa at 1080 C and then decreases to 927 MPa at 1120 C which is close related to the grain growth of welding nugget zone (WNZ). The change tendency of fatigue limit with STT is in agreement with that of yield strength. It reaches the maximum of 765 MPa at 1080 C which is closely connected to the shape transition of ?? precipitates in WNZ and the grain growth. Combining high temperature tensile properties with fatigue limit, the optimal STT is 1080 C.
Vision-based action recognition of construction workers has attracted increasing attention for its diverse applications. Though state-of-the-art performances have been achieved using spatial-temporal features in previous studies, considerable challenges remain in the context of cluttered and dynamic construction sites. Considering that workers actions are closely related to various construction entities, this paper proposes a novel system on enhancing action recognition using semantic information. A data-driven scene parsing method, named label transfer, is adopted to recognize construction entities in the entire scene. A probabilistic model of actions with context is established. Worker actions are first classified using dense trajectories, and then improved by construction object recognition. The experimental results on a comprehensive dataset show that the proposed system outperforms the baseline algorithm by 10.5%. The paper provides a new solution to integrate semantic information globally, other than conventional object detection, which can only depict local context. The proposed system is especially suitable for construction sites, where semantic information is rich from local objects to global surroundings. As compared to other methods using object detection to integrate context information, it is easy to implement, requiring no tedious training or parameter tuning, and is scalable to the number of recognizable objects.
Background subtraction plays a very important role in video analysis, especially in surveillance systems. While being straightforward, the performance based on frame differencing is unsatisfied due to its sensitiveness to issues such as camera shake and swinging objects. To address its limitations, in this paper we propose a complete adaptive difference modelling framework. First, we introduce two difference discriminators to model the evolution process of pixels. Second, we use Gaussian Mixture Models to adaptively learn the difference threshold to distinguish foreground from background. Third, three heuristics are employed to further improve the model adaptability. Experiments on real-world videos of the Background Models Challenge (BMC) demonstrate that our method performs better on global quality metric (FSD) than other state-of-the-art methods.
High-entropy alloys(HEAs),with a new alloy-ing concept,could possess many unique mechanical and functional properties.The current work investigated whe-ther one such alloy offers potential for bearing surfaces under dry conditions.The dry,reciprocating sliding wear characteristics of AlCoCrFeNiTi0.5 alloy were investigated under various applied loads and sliding speeds.Transmis-sion electron microscopy(TEM)and scanning electron microscopy(SEM)were utilized to characterize internal structure and wear surfaces of the alloy,respectively.It is found that the AlCoCrFeNiTi0.5 alloy preserves better wear resistance than Fe77Ni23 solid solution alloy,Ti-46Al-2Cr-2Nb intermetallic alloy,or a wear-resistant steel AISI 52100,especially under higher loads.The wear rate increases slowly with the applied loads increasing and keeps steady under different sliding speeds.The wear mechanisms are abrasive wear,adhesive wear and oxida-tive wear.The nano-sized Fe-Cr solid solution and Al-Ni-Ti rich intermetallic phase precipitated in the dendritic regions and the formation of oxidation play important roles in the good wear resistances of this high-entropy alloy.
Joint Segmentation and Recognition of Worker Actions using Semi-Markov Models Jun Yang, Zhongke Shi and Ziyan Wu Pages 522-528 (2016 Proceedings of the 33rd ISARC, Auburn, USA, ISBN 978-1-5108-2992-3, ISSN 2413-5844) Abstract: Vision-based automated recognition of worker actions has gain lots of interest during the past few years. However, existing research all requires pre-segmented video clips, which is not applicable in the real situation. Furthermore, pre-segmented videos abandon the temporal information of action transition. A joint action segmentation and recognition method, which can segment continuous video stream while recognizing the action type for each segment, is an urgent need. In this paper, we model the worker actions with a discriminative semi-Markov model. In the model, a set of features is defined to capture both the local and global characteristics of each action cycle. Then the semi-Markov model is formulated as an optimization problem and solved by the cutting plane method for simultaneous action segmentation and recognition. Scale-Invariant Feature Transform (SIFT) is applied to detect feature points in the region of interest in every frame. Two descriptors (Histograms of Oriented Gradients ? HOG, Histograms of Optical Flow ? HOF), are computed in the feature points to encode the scenario and motion flow simultaneously. Finally, the Bag-of-Feature strategy is adopted for feature representation. Experimental results from real world construction videos show that the proposed method is able to segment and recognize continuous worker actions correctly, resulting in a prospecting application in automated productivity analysis. Keywords: Action Segmentation, Action Recognition, Worker, Semi-Markov Models DOI: https://doi.org/10.22260/ISARC2016/0063 Download fulltext Download BibTex Download Endnote (RIS) TeX Import to Mendeley
The performance of an object detection system relies heavily on two components: an object model to capture the compositional relationship among the object body and its parts, and a feature representation to describe object appearance. In this work, we present an empirical study of combining two state-of-the-art such components: Deformable Part Model (DPM), a proven effective and flexible part-based object model which originally adopts Histogram of Oriented Gradients (HOG) feature, and Aggregated Channel Features (ACF), a unified feature representation framework with fast pyramid calculation which is originally used in a rigid template matching scheme. DPM is known to work but slow, at the same time ACF has previously been shown to yield a massive speedup with only a minor loss in accuracy compared to competing features including HOG. By combining the two, our hope is to achieve the best of both worlds: the object structure representation power of DPM and the computational efficiency of ACF. Our experiments show that while ACF with heterogeneous feature channels could improve the accuracy of DPM, the run time benefit introduced by fast pyramid approximation is rather limited.
Wide spread monitoring cameras on construction sites provide large amount of information for construction management. The emerging of computer vision and machine learning technologies enables automated recognition of construction activities from videos. As the executors of construction, the activities of construction workers have strong impact on productivity and progress. Compared to machine work, manual work is more subjective and may differ largely in operation flow and productivity among different individuals. Hence only a handful of work studies on vision based action recognition of construction workers. Lacking of publicly available datasets is one of the main reasons that currently hinder advancement. The paper studies worker actions comprehensively, abstracts 11 common types of actions from 5 kinds of trades and establishes a new real world video dataset with 1176 instances. For action recognition, a cutting-edge video description method, dense trajectories, has been applied. Support vector machines are integrated with a bag-of-features pipeline for action learning and classification. Performances on multiple types of descriptors (Histograms of Oriented Gradients – HOG, Histograms of Optical Flow – HOF, Motion Boundary Histogram – MBH) and their combination have been evaluated. Discussion on different parameter settings and comparison to the state-of-the-art method are provided. Experimental results show that the system with codebook size 500 and MBH descriptor has achieved an average accuracy of 59% for worker action recognition, outperforming the state-of-the-art result by 24%.
As-built building information model (BIM) is an urgent need of the architecture, engineering, construction and facilities management (AEC/FM) community. However, its creation procedure is still lab...
As-built building information model (BIM) is an urgent need of the architecture, engineering, construction and facilities management (AEC/FM) community. However, its creation procedure is still labor-intensive and far from maturity. Taking advantage of prevalence of digital cameras and the development of advanced computer vision technology, the paper proposes to reconstruct a building facade and recognize its surface materials from images taken from various points of view. These can serve as initial steps towards automatic generation of as-built BIM. Specifically, 3D point clouds are generated from multiple images using structure from motion method and then segmented into planar components, which are further recognized as different structural components through knowledge based reasoning. Windows are detected through a multilayered complementary strategy by combining detection results from every semantic layer. A novel machine learning based 3D material recognition strategy is presented. Binary classifiers are trained through support vector machines. Material type at a given 3D location is predicted by all its corresponding 2D feature points. Experimental results from three existing buildings validate the proposed system.
Timely and accurate monitoring of onsite construction operations can bring an immediate awareness on project specific issues. It provides practitioners with the information they need to easily and quickly make project control decisions. Despite their importance, the current practices are still time-consuming, costly, and prone to errors. To facilitate the process of collecting and analyzing performance data, researchers have focused on devising methods that can semi-automatically or automatically assess ongoing operations both at project level and operation level. A major line of work has particularly focused on developing computer vision techniques that can leverage still images, time-lapse photos and video streams for documenting the work in progress. To this end, this paper extensively reviews these state-of-the-art vision-based construction performance monitoring methods. Based on the level of information perceived and the types of output, these methods are mainly divided into two categories (namely project level: visual monitoring of civil infrastructure or building elements vs. operation level: visual monitoring of construction equipment and workers). The underlying formulations and assumptions used in these methods are discussed in detail. Finally the gaps in knowledge that need to be addressed in future research are identified.
Automatic Recognition of Construction Worker Activities Using Dense Trajectories Jun Yang, Zhongke Shi, Ziyan Wu Pages 1-7 (2015 Proceedings of the 32nd ISARC, Oulu, Finland, ISBN 978-951-758-597-2, ISSN 2413-5844) Abstract: Wide spread monitoring cameras on construction sites provide large amount of information for construction management. The emerging of computer vision and machine learning technologies enables automatic recognition of construction activities from videos. As the executors of construction, the activities of construction workers have strong impact on productivity and progress. Compared to machine work, manual work is more subjective and may differ largely in operation flow and productivity from one worker to another. Hence only a handful of work study on vision based activity recognition of construction workers. Lacking of publicly available datasets is one of the main reasons that currently hinder advancement. The paper studies manual work of construction workers comprehensively, selects 11 common types of activities and establishes a new real world video dataset with 1176 instances. For activity recognition, a cutting-edge video description method, dense trajectories, has been applied. Support vector machines are integrated with a bag-of-features pipeline for activity learning and classification. Performance on multiple types of descriptors (Histograms of Oriented Gradients - HOG, Histograms of Optical Flow - HOF, Motion Boundary Histogram - MBH) and their combination has been evaluated. Experimental results show that the proposed system has achieved a state-of-art performance on the new dataset. Keywords: Worker, Activity recognition, Dense trajectory, Machine learning DOI: https://doi.org/10.22260/ISARC2015/0007 Download fulltext Download BibTex Download Endnote (RIS) TeX Import to Mendeley
3D building modeling has many potential uses in the fields of construction, city planning and public security. An image-based 3D semantic modeling method of building facade is proposed in this paper. Dense point clouds are generated from inputting images by structure from motion and cluster based multi-view-stereo algorithms. Planar components are extracted from generated point clouds by random sample consensus and further recognized as structural components based on prior knowledge. Windows are detected through a multi-layer complementary strategy with binary image processing techniques. Experimental results from two building facades verify the proposed method.
Information technology has been applied in construction management and is gaining in popularity in recent years. This paper explored the state-of-the-art research and made a comprehensive literature review from two aspects: ‘IT supported construction process surveillance’ and ‘IT supported built infrastructure evaluation’. Multiple sensor technologies (including GPS, RFID, UWB, video camera and laser scanner) and their applications in construction management were introduced separately. Then some of the author’s ongoing research was introduced briefly with two future perspectives envisaged. In the near future IT supported construction management is expected to become a rapidly increasing area and has far-reaching impact in industry.