Temporal Action Detection (TAD) plays a vital role in video understanding by precisely determining what happens and where it occurs in untrimmed streams. While Transformers have improved TAD performance, existing methods struggle to extract sufficiently discriminative representations. Furthermore, they employ separate heads for classification and localization but share the same input features, making it difficult to balance the two tasks. We introduce E-FFT-DE to overcome these bottlenecks. Specifically, an Efficient FFT-based Encoder (E-FFT) captures global temporal contexts to build highly representative features. Coupled with this, a Decoupled Enhancement (DE) module fuses multi-scale semantics to generate task-specific features, effectively eliminating the task trade-off. Combined, our E-FFT-DE model achieves promising performance on THUMOS14, MultiTHUMOS, ActivityNet-1.3, and HACS.
The Wide-field X-ray Telescope (WXT) onboard the Einstein Probe (EP) produces a large post-detection candidate stream in which genuine astrophysical sources coexist with instrumental artifacts and Cosmic Ray events. We present M-EPDet, a three-step post-detection framework for real-time candidate vetting in EP-WXT lobster-eye Micro-pore Optics (MPO) data. The framework combines a ResNet-based Arm filter, a dual-branch temporal-spectral Cosmic Ray filter, and a background-aware Bayesian Blocks module for single-exposure variability screening. Using on-orbit EP-WXT observations, we report decoupled metrics for the cascading system. M-EPDet achieves a Real-Bogus Recall of 98.31% (98.53%× 99.78%) for genuine astrophysical sources, together with rejection rates of 92.99% for instrumental artifacts and 98.18% for Cosmic Ray events. In the final step, the Bayesian Blocks module flags 0.75% of the post-filtration observations, corresponding to a 99.25% reduction in candidate volume. The system is deployed in the EP-WXT pipeline as a lightweight real-time service, reducing the manual-inspection burden in candidate vetting.
Temporal Action Localization (TAL) requires jointly optimizing action classification and temporal boundary regression. However, these two tasks favor conflicting feature representations: classification relies on semantically rich context, whereas localization demands fine-grained temporal details. Most existing one-stage TAL detectors decouple these tasks only at the prediction heads while sharing intermediate features, causing task interference and limiting boundary precision. To address this, we propose Text-Guided Decoupling for Temporal Action Localization (TGD-TAL), a text-conditioned framework that disentangles task-specific representations using video-derived textual priors. Specifically, we employ a Multimodal Large Language Model (MLLM) to generate descriptive text, which is encoded as semantic priors to guide feature modulation via a Text-Conditioned Feature Modulation (TCFM) module. Building upon a temporal feature pyramid, we design two complementary branches: the Text-Guided Classification Decoupling (TGCD) enhances category discrimination by attending to high-level contextual features, while the Boundary-Aware Text-Guided Regressor (BTR) refines temporal boundaries by leveraging low-level details and start- and end-aware textual cues via text–video cross-attention and gated fusion. Experiments show our method outperforms baselines, validating text priors’ effectiveness in reducing interference and sharpening boundaries.
Accurate prediction of fatigue crack growth rate (FCGR) is of great significance in the field of materials science and engineering. To address the limitations of existing prediction methods due to data scarcity and suboptimal model performance, this paper proposes a method called the Data-Augmented and Dual-Optimized Neural Network (DA-DONN). The method employs piecewise cubic Hermite interpolation (PCHIP) to augment the FCGR data, thereby alleviating the small-sample problem. Bayesian optimization (BO) is used to tune the hyperparameters of the neural network, and the horned lizard optimization algorithm (HLOA) is applied to optimize the initial weights and biases, thus improving the prediction accuracy. Experimental results indicate that DADONN achieves a significantly lower MSE than SA-DNN on the 7075 aluminum alloy test set. On the 6013 aluminum alloy dataset, DA-DONN also outperforms DLFCO-DNN and MFA-DNN in terms of MAE, RMSE, and Mean RE, demonstrating its superior accuracy and practical feasibility.
Failure in engineering structures is predominantly caused by the propagation of Mode I-II fatigue cracks. To address the limitations of traditional methods and machine learning approaches, this study proposes a hybrid prediction method that combines the Finite Element Method (FEM) with optimized neural networks for the prediction of Mode I-II fatigue crack growth (FCG) paths. Initially, the GOOSE Optimization Algorithm (GOOSE) is employed to optimize the initial weights and biases of the neural network, thereby enhancing its predictive accuracy. The optimized neural network is then used to predict the crack growth paths, and at each prediction step, it dynamically evaluates whether FEM analysis is required for error correction. Experimental results indicate that the proposed method significantly reduces the Mean Squared Error (MSE) in crack growth path predictions when compared to pure machine learning approaches. Moreover, in comparison with pure FEM, this method significantly enhances computational efficiency while maintaining prediction accuracy, thus demonstrating its performance and advantages.
Seismic traveltime is a crucial seismic attribute that directly influences the computational accuracy and efficiency of various seismic processing and interpretation methods. The Fast Marching Method (FMM) is an accurate and stable finite-difference approach for calculating seismic traveltime. However, when applied to three-dimensional models, its computational efficiency becomes a major limitation. To address this issue, this paper proposes an efficient traveltime computation algorithm that integrates coarse-grid interpolation with the FMM. The proposed method first computes the global traveltime field on a coarse grid using FMM, then refines the results to a fine grid across the entire domain via cubic spline interpolation. Both theoretical analysis and numerical experiments demonstrate that the proposed approach preserves the unconditional stability of the FMM while significantly reducing memory consumption and computational time by combining coarse-grid computation with interpolation-based refinement.
Deep learning (DL) has demonstrated significant potential for direct seismic impedance inversion in the depth domain. However, conventional DL frameworks often adopt a trace-by-trace strategy, which fails to account for the spatial coherence of subsurface structures, leading to lateral discontinuities and a lack of geological guidance. To address these limitations, we propose a dip-constrained multitrace deep neural network (DMDNN) for depth-domain seismic inversion. This approach integrates structural dip as a geological prior through a sliding window mechanism to quantify the spatial continuity of strata, thereby providing the explicit geological guidance missing in traditional methods. Unlike the standard multitrace input paradigm, the proposed method establishes physical correlations between adjacent traces via an implicit lateral constraint mechanism in the loss function, enabling effective multitrace collaborative inversion even with single-trace inputs. Furthermore, a Huber loss function is employed for label constraints to enhance the framework's resilience to seismic noise, offering superior stability over standard mean squared error (mse) metrics. Extensive evaluations on synthetic and field datasets demonstrate that, with properly calibrated weights for the structural dip loss and constraint window sizes, the DMDNN significantly enhances lateral continuity and vertical resolution while maintaining high geological plausibility. By reducing dependence on large-scale training samples, this method provides a robust and geologically consistent pathway for reservoir characterization in complex structural zones.
Pre-stack seismic inversion is used to calculate elastic parameters, including P-wave and S-wave velocities, as well as densities. These parameters play an integral role in the characterization of reservoirs, thereby enhancing the exploration and production process. Deep learning-based seismic inversion does not need a known physical system and can give satisfactory results with sufficient training data. The acquisition of such datasets for seismic inversion poses a significant challenge due to the exorbitant costs associated with drilling activities. Integrating domain knowledge, physical systems, and well log data into a deep learning-based seismic inversion framework is crucial for improving its efficiency and effectiveness. Nevertheless, existing data-driven approaches do not adequately exploit such information, thereby constraining their overall performance and applicability. Therefore, we develop a double dual neural network structure built upon the closed-loop neural network framework, which incorporates both physics and model information to mitigate the dependency on extensive labeled datasets. The information from the different domains is linked through a loss function, where one dual network is responsible for constraining the inversion results using physics information to ensure the physics consistency of the predictions, and the other dual network is responsible for constraining the inversion results using a priori model information to enhance the reliability of the predictions. The method makes full use of well-log data for network training when wells are available, as well as providing unsupervised learning and inversion under well-free conditions. The integration of qualitative and quantitative analyses proves instrumental in demonstrating the effectiveness of the proposed methodology through the use of synthetic and field pre-stack examples.
The quasi-zero stiffness metastructures have excellent low-frequency vibration isolation performance. However, due to its unique working principle, it can only effectively isolate objects subjected to specific loads and is greatly limited in many application scenarios. A quasi-zero-stiffness metastructure with a distinct structural configuration from previous designs was proposed, and the optimization of its quasi-zero-stiffness characteristics was achieved using multi-objective optimization theory. The mechanical properties and vibration isolation performance are comprehensively analyzed through theoretical analysis, finite element simulation, and experimental testing. The structure customization theory and design methodology for multi-stage quasi-zero-stiffness systems are summarized and validated via finite element simulations. It has been proved that the proposed quasi-zero stiffness structure has excellent low-frequency vibration isolation performance. The proposed structural customization theory can precisely customize a quasi-zero stiffness lattice with specific loads based on structural parameters. The multi-level quasi-zero stiffness design method utilizes various materials and combinations of different lattice layers to obtain a quasi-zero stiffness system with multiple vibration isolation platforms. The proposed double-cosine beam quasi-zero stiffness vibration isolation structure in this paper provides certain supplementation to the design concept of metastructures. The customized structure idea enhances the accuracy of its vibration isolation. The design of multi-level quasi-zero stiffness vibration isolation systems strengthens the practicability of the quasi-zero stiffness system.
YOLOv10, known for its efficiency in object detection methods, quickly and accurately detects objects in images. However, when detecting small objects in remote sensing imagery, traditional algorithms often encounter challenges like background noise, missing information, and complex multiobject interactions, which can affect detection performance. To address these issues, we propose an enhanced algorithm for detecting small objects, named SOD-YOLOv10. We design the Multidimensional Information Interaction for the Transformer Backbone (TransBone) Network, which enhances global perception capabilities and effectively integrates both local and global information, thereby improving the detection of small object features. We also propose a feature fusion technology using an attention mechanism, called aggregated attention in a gated feature pyramid network (AA-GFPN). This technology uses an efficient feature aggregation network and re-parameterization techniques to optimize information interaction between feature maps of different scales. Additionally, by incorporating the aggregated attention (AA) mechanism, it accurately identifies essential features of small objects. Moreover, we propose the adaptive focal powerful IoU (AFP-IoU) loss function, which not only prevents excessive expansion of the anchor box area but also significantly accelerates model convergence. To evaluate our method, we conduct thorough tests on the RSOD, NWPU VHR-10, VisDrone2019, and AI-TOD datasets. The findings indicate that our SOD-YOLOv10 model attains 95.90%, 92.46%, 55.61%, and 59.47% for mAP@0.5 and 73.42%, 66.84%, 39.03%, and 42.67% for mAP@0.5:0.95.
As an important seismic attribute parameter, accurate travel-time computation is essential for simulating seismic wave propagation, locating seismic sources, constructing velocity models, and enabling earthquake early warning systems. In practical scenarios, the influence of complex topography often necessitates the use of nonhorizontal acquisition lines for seismic surveys. Therefore, investigating seismic travel-time computation under undulating terrain is of significant importance. This study systematically investigates travel-time computation methodologies utilizing coarse-grid interpolation techniques in irregular topographic environments and subsequently proposes an enhanced algorithm. Extending prior accuracy, the proposed method adds new grid node types. It adopts a hybrid interpolation strategy, effectively addressing the inefficiency of traditional methods when applied to large-scale models. This enhancement offers a reliable and efficient solution for computing travel times in large-scale, complex velocity models with topographic variations. The effectiveness of the proposed method is validated through error analysis and multiple numerical experiments.
Denoising of seismic signals is a critical preprocessing step in seismic data analysis, directly impacting the accuracy of subsequent data processing. To enhance the effectiveness of seismic signal denoising, this study proposes a method that combines Variational Mode Decomposition (VMD) optimized by Particle Swarm Optimization (PSO) with an improved Wavelet Transform (WT). First, the key parameters (number of modal functions and penalty factor) of VMD are adaptively selected using an enhanced PSO algorithm to ensure effective decomposition and mitigate mode mixing. Then, for each Intrinsic Mode Function (IMF) obtained from VMD, an improved wavelet denoising technique incorporating a modified thresholding function is applied to further suppress high-frequency noise while preserving signal details. Comparative experiments on both synthetic and field seismic data demonstrate that the proposed method outperforms traditional VMD in terms of signal-to-noise ratio (SNR) improvement and waveform preservation. These results highlight the method’s practicality and robustness in complex signal environments, offering strong support for seismic data preprocessing and interpretation.
This paper presents a temporal action detection (TAD) method with multi-granularity feature aggregation and cross-level boundary modeling (MGCBM). Compared with other methods, our proposed approach has the following advantages. First, different from most existing works which only consider the local temporal context, a simple and computationally efficient Multi-Granularity (MG) module is proposed to comprehensively extract video features in instant, local and global temporal granularities. Second, unlike the methods that only employ the information from single feature pyramid level for action boundary regression, a cross-level boundary modeling (CBM) strategy that integrates the relative information from both the same and higher level features is designed to improve the accuracy of boundary prediction. At last, benefiting from the MG module and CBM strategy, our method outperforms other state-of-the-art approaches on five challenging TAD datasets: THUMOS14, MultiTHUMOS, EPIC-KITCHENS-100, ActivityNet-1.3 and HACS. We make our code and pre-trained model publicly available at: https://github.com/MGCBM/TAL-MGCBM.
Based on the observational data from 60 short-period stations deployed in the Jishishan M6.2 earthquake epicenter and adjacent regions (Gansu Province, 2023), this study inverted the near-surface S-wave velocity structure through teleseismic receiver function analysis by using the amplitude of direct P-wave. The results reveal that the epicentral area (Liugou Township and surroundings) exhibits markedly low S-wave velocities of 400-600 m/s, with a mean value of (500 ± 50) m/s. In contrast, intermountain basins—Guanting Basin and Dahejia Basin—demonstrate significantly elevated velocities, exceeding the epicentral zone by 100–300 m/s, with values concentrated at 600–900 m/s. Notably, localized areas such as Jintian Village and Caotan Village maintain stable S-wave velocities of (700 ± 30) m/s.The western margin tectonic belt of Jishishan displays distinctive velocity diff erentiation: A pronounced velocity gradient zone along the 35.8°N latitude boundary separates northern areas (<550 m/s) from southern regions (>750 m/s). These findings demonstrate significant spatial heterogeneity in shallow S-wave velocity structures, primarily controlled by three factors: (1) topographic-geomorphic units, (2) stratigraphic lithological contrasts, and (3) anthropogenic modifications. The persistent low-velocity anomalies (<600 m/s) in the epicentral zone and northern Yellow River T2 terrace likely correlate with Quaternary unconsolidated sediments, enhanced groundwater circulation, and bedrock weathering.These results provide critical geophysical constraints for understanding both the seismogenic environment of the Jishishan earthquake and its damage distribution patterns. Furthermore, they establish a foundational framework for regional seismic intensity evaluation, site amplification analysis, and secondary hazard risk assessment.
Computing seismic traveltime constitutes an essential task for seismic imaging and inversion processes. Accurate traveltime computation provides essential parameters for simulating seismic wave propagation, locating seismic sources, establishing velocity models, and seismic imaging. Therefore, a method for calculating seismic traveltime that is accurate, efficient, and stable holds significant research and practical value. The Multi-Stencils Fast Marching (MSFM) method improves upon the Fast Marching Method by employing multiple stencils to enhance accuracy, addressing the issue of significant errors in diagonal directions. However, its computational complexity increases exponentially with the number of stencils, leading to efficiency bottlenecks in large-scale models. This study proposes an enhanced MSFM method incorporating coarse-grid interpolation to compute seismic traveltime in complex velocity models. The technique extends the conventional MSFM framework by refining its narrowband expansion mechanism, utilizing coarse-grid technology to interpolate values from coarse-grid nodes to fine-grid nodes, thereby improving computational efficiency and reducing memory usage. The improved MSFM algorithm is validated using multiple models, demonstrating that it significantly enhances computational efficiency while maintaining good accuracy.
Belt conveyors are one of the key transportation equipment for continuously conveying bulk materials in bulk cargo terminals, and their complex operational environments pose challenges for real-time visual detection of conveyor belt damage. In light of the severe degradation of image feature information concerning longitudinal tears on conveyor belts in complex environments, which increases the difficulty of extracting damage features, a conveyor belt damage detection method based on the YOLOv8 visual detection model is proposed. Firstly, by replacing CSPDarknet53 with the new lightweight network RepVit as the backbone feature extraction network and optimising MobileNet-V3-L based on the architecture of Vision Transformers, the enhancement of image feature extraction capability under complex backgrounds is achieved. Subsequently, a deformable attention mechanism (DAT) is integrated at the back end of the RepVit network to focus the network more on the damage target detection part, adaptively adjusting the receptive fields to better handle variations in damage feature shapes and sizes, enabling rapid and accurate detection of surface damage on conveyor belts. Furthermore, lightweight processing is applied to the YOLOv8 detection head, altering the convolution method to decrease network computational complexity and parameters, thereby enhancing network detection accuracy. The experimental results show that the proposed method is highly reliable, effective, and performs in real time. The model achieves a mean average precision of 93.70% on a dataset of complex operating conditions, demonstrating a significant improvement in detection performance compared to other networks. The frames per second is 25.8, meeting the requirements for real-time detection of conveyor belt damage.
Vision-language pretraining (VLP) models have demonstrated exceptional performance across a wide range of image-text multimodal tasks. Despite their prominence, research confirms that these systems retain significant susceptibility to adversarial manipulation. Existing multimodal adversarial attack methods often fail to fully exploit sample-specific semantic structures, resulting in suboptimal cross-modal alignment and limited transferability of adversarial examples. To overcome this limitation, we propose MGSA-a Multi-Granularity Semantic Alignment Attack framework that enhances adversarial perturbation transferability by jointly disrupting cross-modal semantics at both global and fine-grained levels. MGSA captures coarse-grained alignment using overall representations and fine-grained correspondence by selectively aggregating key image regions and words based on importance. This dual-level joint optimization effectively perturbs both holistic consistency and detailed correspondences, thereby significantly enhancing attack effectiveness in white-box scenarios and transferability to black-box models. Extensive experiments conducted across diverse model architectures and multimodal tasks demonstrate that our method achieves strong performance in white-box settings while significantly improving black-box attack success rates. The results highlight the vulnerability of current VLP models and the effectiveness of our approach in generating transferable and semantically grounded adversarial examples.
In seismic exploration, converted wave data contains important stratigraphic and structural information. Common conversion point (CCP) gather extraction is a key step in converted wave seismic data processing, which has a great impact on velocity modeling and migration imaging. The traditional CCP selection and sorting extraction method mainly adopts the whole trace selection and sorting method, typically the asymptotic conversion point (ACP) gather extraction. This method is able to better meet the imaging requirement of deep strata through the classification strategy of asymptotic approximation, but it cannot image the shallow strata well. To solve the above problems, a new spatiotemporally variable conversion point gather extraction method is proposed. The method gives the mapping relationship between the input converted wave data and the output CCP gather by accurately calculating the spatial locations of the CCP and combining the converted wave two-way traveltime. This strategy allows the reflected seismic signals to be mapped to more accurate spatialtemporal locations, which can effectively deal with the variations of converted waves in asymmetric propagation paths and significantly improve the quality of the CCP gather extraction. The verification results of numerical simulation and field data show that the CCP gather extraction method proposed in this article has demonstrated good application effects in application. The new method not only improves the utilization of data, but also enhances the imaging accuracy of shallow layers and geological structures in complex areas. The new method can be treated as an effective gather extraction tool for seismic converted wave data processing.
Seismic data reconstruction is a critical step in seismic exploration, for which the high-resolution hyperbolic Radon transform (HRT) serves as a key processing tool. This study addresses the challenge of preserving Amplitude Variation with Offset (AVO) characteristics during high-resolution hyperbolic Radon transforms for seismic data reconstruction. Conventional sparse Radon transforms enhance resolution but degrade AVO information due to offset averaging. To overcome this limitation, we propose a novel framework integrating: Fast Iterative Shrinkage Thresholding Algorithm (FISTA) to solve the hyperbolic Radon transform with mixed-norm regularization, significantly improving computational efficiency and Radon panel resolution compared to the Conjugate Gradient (CG) method. Polynomial fitting of AVO attributes within the inverse Radon transform, enabling accurate recovery of amplitude-offset relationships. Synthetic data and field data demonstrates that FISTA achieves superior sparsity in the Radon domain. The polynomial-augmented inverse transform preserves AVO features and reduces reconstruction residuals versus conventional methods. This method provides robust technical support for the accurate interpretation of seismic data and for subsequent hydrocarbon exploration. Consequently, it holds significant practical value and shows extensive prospects for future application.