Video understanding technology has become increasingly important in various disciplines, yet current approaches have primarily focused on lower comprehension level of video content, posing challenges for providing comprehensive and professional insights at a higher comprehension level. Video analysis plays a crucial role in athlete training and strategy development in racket sports. This study aims to demonstrate an innovative and higher-level video comprehension framework (ChatMatch), which integrates computer vision technologies with the cutting-edge large language models (LLM) to enable intelligent analysis and inference of racket sports videos. To examine the feasibility of this framework, we deployed a prototype of ChatMatch in the badminton in this study. A vision-based encoder was first proposed to extract the meta-features included the locations, actions, gestures, and action results of players in each frame of racket match videos, followed by a rule-based decoding method to transform the extracted information in both structured knowledge and unstructured knowledge. A set of LLM-based agents included namely task identifier, coach agent, statistician agent, and video manager, was developed through a prompt engineering and driven by an automated mechanism. The automatic collaborative interaction among the agents enabled the provision of a comprehensive response to professional inquiries from users. The validation findings showed that our vision models had excellent performances in meta-feature extraction, achieving a location identification accuracy of 0.991, an action recognition accuracy of 0.902, and a gesture recognition accuracy of 0.950. Additionally, a total of 100 questions were gathered from four proficient badminton players and one coach to evaluate the performance of the LLM-based agents, and the outcomes obtained from ChatMatch exhibited commendable results across general inquiries, statistical queries, and video retrieval tasks. These findings highlight the potential of using this approach that can offer valuable insights for athletes and coaches while significantly improve the efficiency of sports video analysis.
Characterizing multivariate parameters is crucial for uncertainty analysis in geological engineering. However, the commonly used multivariate distribution function-based methods are increasingly challenged by issues of diminished accuracy and efficiency, as well as the inclusion of subjective processes, due to the growing complexity of data structures resulting from advanced data acquisition techniques. To this end, the prevailing generative machine learning algorithms are tried to learn the joint distributions of multivariate geotechnical parameters, and a deep neural network, CasMDN, is proposed based on the mixture density network and a cascade mechanism. The principle, structure, and training method of CasMDN is first introduced. Then two applications of the model, including stochastic simulation and probability calculation, are presented. In the first experiment, a group of geometric parameters of rock joints are simulated by a variational autoencoder, a generative adversarial network, a Gaussian mixture model, and our model, respectively. The result shows a significant advantage of our method over other generative models in learning the joint distribution of orientation, trace length, and aperture. Then, another comparison between our method and the recently proposed copula-based approach is carried out with a set of five-dimensional soil parameters, which is an open-source dataset collected from previous literature. It is proved that the CasMDN model can achieve the best performances on both fitting and probability calculating. This research demonstrates that using intelligent generative models to characterize complex geotechnical parameters can help explore the deep laws of geological uncertainty by improving accuracy and reducing subjectivity. Furthermore, the proposed CasMDN model is considered to have a broad application scope beyond the provided cases, for instance, uncertainty modelling and risk assessment of geotechnical engineering.
Digital construction relies on effective sensing to enhance the safety, productivity, and quality of its activities. However, current sensing devices (e.g., camera, LiDAR, infrared sensors) have significant limitations in different aspects. In light of the substantial advantages offered by emerging 4D mmw technology, it is believed that this technology can overcome these limitations and serve as an excellent complement to current construction sensing methods due to its robust imaging capabilities, spatial sensing abilities, velocity measurement accuracy, penetrability features, and weather resistance properties. To support this argument, a scientometric review of 4D mmw-based sensing is conducted in this study. A total of 213 articles published after the initial invention of 4D mmw technology in 2019 were retrieved from the Scopus database, and six kinds of metadata were extracted from them, including the title, abstract, keywords, author(s), publisher, and year. Since some papers lack keywords, the GPT-4 model was used to extract them from the titles and abstracts of these publications. The preprocessed metadata were then integrated using Python and fed into the Citespace 6.2.R3 for further statistical, clustering, and co-occurrence analyses. The result revealed that the primary applications of 4D mmw are autonomous driving, human activity recognition, and robotics. Subsequently, the potential applications of this technology in the construction industry are explored, including construction site monitoring, environment understanding, and worker health monitoring. Finally, the challenges of adopting this emerging technology in the construction industry are also discussed.
Limited by the survey data and current interpretation methods, the modelling processes of fault networks are fraught with uncertainties. In hydraulic geological engineering, the location uncertainty of faults plays a vital role in decision-making and engineering safety. However, traditional uncertainty modelling methods have difficulty obtaining accurate uncertainty quantification and topology representation. To this end, we proposed a novel solution for uncertainty analysis and three-dimensional modelling for faults via a deep learning approach. A spatial uncertainty perception (SUP) method is first presented based on a modified deep mixture density network (MDN), which can be used to learn the spatial distributions of fault zones, calculate the probability of fault models, and simulate stochastic models with certain confidence degrees. After that, a graph representation (GRep) method is developed to express the topological form and geological ages of fault networks. The GRep makes it possible to automatically simulate the spatial distributions of fault belts, thus providing an effective way for the uncertainty modelling and assessment of fault networks. The two methods are then performed in the geological engineering of a practical hydraulic project. The results show that this solution can conduct accurate uncertainty evaluations and visualizations on fault networks, thus providing suggestions for subsequent geological investigations.
Seismic performance of gravity dams has been widely concerned due to their importance in hydraulic engineering. In recent years, the influence of the incident angle of seismic waves and the inhomogeneity of complex foundation on the soil-structure interaction (SSI) system has become a focus. However, the foundation is often simplified to homogeneous elastomer when simulating nonuniform excitation seismic wave propagation due to the complex theoretical formulations. Therefore, this paper considers the stratification of foundation and analyzes the dynamic response of a gravity dam under obliquely incident seismic waves. First, the gravity dam-layered foundation interaction system is established in the finite-element software, and the concrete damage plasticity (CDP) model is adopted of the dam. Then, the displacements of the foundation under seismic P or SV waves with arbitrary incident angles are calculated by combining the one-dimensional time-domain method and the free wave field calculation method. Finally, adopting the wave input method based on the substructure of artificial boundaries, the seismic waves are converted to equivalent nodal forces on the viscous-spring boundaries of the layered foundation. Furthermore, the dynamic response of the dam on the homogeneous foundation is also simulated for comparison in this paper. The calculation results confirm that the impact of arbitrary incident angle of earthquakes and the parameters of foundation on the dam response is significant, and both should be considered comprehensively.
The determination of the incident angle of an earthquake is one of the critical research issues in modeling the seismic input mechanism at a dam site and it is also for geological exploration. At present, the incident angle calculation methods mostly rely on accurate geological models, complex theoretical formulations, and complicated procedures. To this end, an incident angle estimation method using dam vibration response data and a machine learning algorithm is presented in this study. First, a three-dimensional finite element gravity dam-foundation system model is constructed. A wave input method based on the viscous-spring artificial boundary is used to simulate the multi-angle incidence of P and SV waves. The vibration responses of key dam points are obtained. Nine features that have geometric interpretation are constructed by the response data and the new data are substituted into a stacking ensemble algorithm for training. Finally, the angles of obliquely incident P and SV waves are estimated using the trained stacking ensemble estimation model. The results reveal correlation between dam responses and the incident angle with the average R 2 values of the estimation model of 0.996 (P waves) and 0.996 (SV waves), and the average root mean square errors of 1.765° (P waves) and 0.546° (SV waves). It is thus confirmed that the estimation model integrated multiple features from several measurement points with high accuracy and stability. In addition, the proposed method can be extended to other types of large structures because of its universality.