Parameter-efficient tuning (PEFT) techniques like low-rank adaptation (LoRA) offer training efficiency on Large Language Models, but their impact on model performance remains limited. Recent efforts integrate LoRA and Mixture-of-Experts (MoE) to improve the performance of PEFT methods. Despite promising results, research on improving the efficiency of LoRA with MoE is still in its early stages. Recent studies have shown that experts in the MoE architecture have different strengths and also exhibit some redundancy. Does this statement also apply to parameter-efficient MoE? In this paper, we introduce a novel parameter-efficient MoE method, MoE-LoRA with Layer-wise Expert Allocation (MoLA) for Transformer-based models, where each model layer has the flexibility to employ a varying number of LoRA experts. We investigate several architectures with varying layer-wise expert configurations. Experiments on six well-known NLP and commonsense QA benchmarks demonstrate that MoLA achieves equal or superior performance compared to all baselines. We find that allocating more LoRA experts to higher layers further enhances the effectiveness of models with a certain number of experts in total. With much fewer parameters, this allocation strategy outperforms the setting with the same number of experts in every layer. This work can be widely used as a plug-and-play parameter-efficient tuning approach for various applications. The code is available at https://github.com/GCYZSL/MoLA.
The success of data mixing augmentations in image classification tasks has been well-received. However, these techniques cannot be readily applied to object detection due to challenges such as spatial misalignment, foreground/background distinction, and plurality of instances. To tackle these issues, we first introduce a novel conceptual framework called Supervision Interpolation (SI), which offers a fresh perspective on interpolation-based augmentations by relaxing and generalizing Mixup. Based on SI, we propose LossMix, a simple yet versatile and effective regularization that enhances the performance and robustness of object detectors and more. Our key insight is that we can effectively regularize the training on mixed data by interpolating their loss errors instead of ground truth labels. Empirical results on the PASCAL VOC and MS COCO datasets demonstrate that LossMix can consistently outperform state-of-the-art methods widely adopted for detection. Furthermore, by jointly leveraging LossMix with unsupervised domain adaptation, we successfully improve existing approaches and set a new state of the art for cross-domain object detection.
Large Vision-Language Models (LVLMs) offer remarkable benefits for a variety of vision-language tasks. However, a challenge hindering their application in real-world scenarios, particularly regarding safety, robustness, and reliability, is their constrained semantic grounding ability, which pertains to connecting language to the physical-world entities or concepts referenced in images. Therefore, a crucial need arises for a comprehensive study to assess the semantic grounding ability of widely used LVLMs. Despite the significance, sufficient investigation in this direction is currently lacking. Our work bridges this gap by designing a pipeline for generating large-scale evaluation datasets covering fine-grained semantic information, such as color, number, material, etc., along with a thorough assessment of seven popular LVLMs' semantic grounding ability. Results highlight prevalent misgrounding across various aspects and degrees. To address this issue, we propose a data-centric enhancement method that aims to improve LVLMs' semantic grounding ability through multimodal instruction tuning on fine-grained conversations. Experiments on enhanced LVLMs demonstrate notable improvements in addressing misgrounding issues.
We present LOWA, a novel method for localizing objects with attributes effectively in the wild. It aims to address the insufficiency of current open-vocabulary object detectors, which are limited by the lack of instance-level attribute classification and rare class names. To train LOWA, we propose a hybrid vision-language training strategy to learn object detection and recognition with class names as well as attribute information. With LOWA, users can not only detect objects with class names, but also able to localize objects by attributes. LOWA is built on top of a two-tower vision-language architecture and consists of a standard vision transformer as the image encoder and a similar transformer as the text encoder. To learn the alignment between visual and text inputs at the instance level, we train LOWA with three training steps: object-level training, attribute-aware learning, and free-text joint training of objects and attributes. This hybrid training strategy first ensures correct object detection, then incorporates instance-level attribute information, and finally balances the object class and attribute sensitivity. We evaluate our model performance of attribute classification and attribute localization on the Open-Vocabulary Attribute Detection (OVAD) benchmark and the Visual Attributes in the Wild (VAW) dataset, and experiments indicate strong zero-shot performance. Ablation studies additionally demonstrate the effectiveness of each training step of our approach.
While Multimodal Large Language Models (MLLMs) are widely used for a variety of vision-language tasks, one observation is that they sometimes misinterpret visual inputs or fail to follow textual instructions even in straightforward cases, leading to irrelevant responses, mistakes, and ungrounded claims. This observation is analogous to a phenomenon in neuropsychology known as Agnosia, an inability to correctly process sensory modalities and recognize things (e.g., objects, colors, relations). In our study, we adapt this similar concept to define"agnosia in MLLMs", and our goal is to comprehensively evaluate and mitigate such agnosia in MLLMs. Inspired by the diagnosis and treatment process in neuropsychology, we propose a novel framework EMMA (Evaluation and Mitigation of Multimodal Agnosia). In EMMA, we develop an evaluation module that automatically creates fine-grained and diverse visual question answering examples to assess the extent of agnosia in MLLMs comprehensively. We also develop a mitigation module to reduce agnosia in MLLMs through multimodal instruction tuning on fine-grained conversations. To verify the effectiveness of our framework, we evaluate and analyze agnosia in seven state-of-the-art MLLMs using 9K test samples. The results reveal that most of them exhibit agnosia across various aspects and degrees. We further develop a fine-grained instruction set and tune MLLMs to mitigate agnosia, which led to notable improvement in accuracy.
We present the modelling, design, and characterization of a novel fiber Bragg grating (FBG) accelerometer based on a bearing. A cantilevered mechanical structure with bearings can effectively decrease energy loss during vibration and enhance sensitivity. The principle of the accelerometer is analysed by a simplified model. The influence of the structural parameters of the proposed sensor on its sensitivity and resonance frequency is further analysed for structural optimization. Meanwhile, its dynamic behaviour and frequency response characteristics are analysed by numerical simulations. In addition, dual-FBG not only enables temperature compensation but also enhances sensitivity. Experimental results indicate that the FBG accelerometer provides a linear response over a broad frequency range from 0.5 Hz to 40 Hz, with a high sensitivity of 575.8 pm/g. This FBG accelerometer with low frequency and high sensitivity provides effective support for the design of sensors of the same type.
为实现高速动车组车轮踏面擦伤长度定量评估,提出了一种自适应连续小波模极大值算法,应用到基于轴箱振动加速度信号的车轮擦伤状态识别与损伤评估.利用连续小波变换系数模极大值对自适应滤波后的轴箱振动加速度信号突变点进行定位,从而建立擦伤车轮脱离钢轨腾空运行时间与擦伤长度及其影响参数的关系模型.仿真实例结果表明,该方法可在车辆高速运行时实现车轮擦伤程度的定量识别.与其它文献识别结果对比分析表明,该方法原理简单、计算结果精度高.
With increases in train speed and traffic density, polygonal wear of railway wheels arises accordingly, induced by the high impacts between wheels and rails, which is mainly related to operation safety and ride comfort of vehicle system. This work evaluates the effect of wheel polygon shape on the dynamic performance of the wheel set through numerical simulations. The finite element model, which includes the wheel set and the slab track, was established using ANSYS software to study the effects of polygonal wear on the dynamic behavior of the railway wheel. In the model, wheel–rail interaction forces caused by polygon wheel shape were solved using Universal Mechanisms of wear and were then entered into the finite element model. Using the simulation model, the influence of the harmonic order and out-of-roundness amplitude of wheel polygon on transient dynamic behaviors of the wheels namely, the displacement, acceleration, and von Misses equivalent stress were investigated. The results indicate that both the maximum dynamic displacement and Von Misses equivalent stress of the wheel plate show proportionality to the OOR amplitude, the harmonic order and the vehicle velocity. Besides, the maximum Von Misses equivalent stress occurs close to the wheel center, whereas the maximum displacement occurs close to the wheel tread. The findings will provide a theoretical basis for on-board detection methods of monitoring wheel polygonal wear.
The invention is applicable to the technical field of wheel scratch measurement, and provides a wheel scratch length measurement method and device and terminal equipment. The method comprises the steps: obtaining an axle box vibration acceleration signal and a noise signal corresponding to a wheel; processing the axle box vibration acceleration signal according to the noise signal to obtain a scratch vibration signal; performing continuous wavelet transform on the scratch vibration signal to obtain a continuous wavelet transform coefficient modulus maximum sequence of the scratch vibration signal; according to the continuous wavelet transform coefficient modulus maximum sequence, obtaining a sudden change time difference of the axle box vibration acceleration signal corresponding to the wheel scratch; and according to the abrupt change time difference, calculating the wheel scratch length. While the workload is reduced, the scratch length of the wheel can be accurately and quantitatively measured, the principle is simple, and the method is suitable for the high-speed running working condition of the vehicle.
为了实现低频振动的高灵敏度测量,设计了一种基于转动支撑梁的新型光纤布拉格光栅加速度计.通过分析其振动模型和MATLAB数值计算,优化了传感器的结构参数,设计传感器理论灵敏度为1 725 pm/g,固有频率为68.4 Hz.同时通过COMSOL模拟分析传感器的动态特性,其模拟结果与理论分析吻合.频响特性和幅值特性实验结果表明光纤光栅加速度计在加速度0~2 g、工作频率0.5~20 Hz的范围内,传感器加速度特性曲线呈现良好线性关系,灵敏度高达1 495.2 pm/g,重复性良好.该传感器结构简单紧凑,轴承结构有效减少悬臂梁振动过程中的弹性能耗,可显著提高其灵敏度,能够实现低频振动信号的探测.
The increasing need for repairs of polygonized wheels on high-speed railways in China is becoming problematic. At high speeds, polygonized wheels cause abnormal vibrations at the wheel-rail interface that can be detected via axle-box accelerations. To investigate the quantitative relationship between axle-box acceleration and wheel polygonization in both the time and frequency domains and under high-speed conditions, a dynamics model was developed to simulate the vehicle-track coupling system and that considers both wheel and track flexibility. The calculated axle-box accelerations were analyzed by using the improved ensemble empirical mode decomposition and Wigner-Ville distribution time-frequency method. The numerical results show that the maximum axle-box accelerations and their frequencies are quantitatively related to the harmonic order and out-of-roundness amplitude of polygonized wheels. In addition, measuring the axle-box acceleration enables both the detection of wheel polygonization and the identification of the degree of damage.
The invention is applicable to the technical field of wheel-rail relation, and provides a wheel polygon trackside detection method based on a piezoelectric acceleration sensor, which comprises the following steps: carrying out simulation analysis on steel rail vibration response characteristics caused by wheel polygon abrasion to obtain steel rail vibration response characteristics under the action of a wheel polygon; according to the vibration response characteristics of the steel rail, determining the measuring point position of the piezoelectric acceleration sensor on the steel rail and establishing a finite element simulation model of the piezoelectric acceleration sensor and carrying out structural optimization design on the piezoelectric acceleration sensor, and mounting the piezoelectric acceleration sensor subjected to structural optimization at a measuring point position to carry out wheel polygon state trackside detection. The feasibility of the piezoelectric acceleration sensor is explored in the early stage through simulation research, it is guaranteed that the piezoelectric acceleration sensor keeps long-time stability in the external environment, and therefore the real-time online monitoring requirement for high-speed railway wheel track damage in China is met, and the method has important reference value for engineering application.