In a complex tracking environment,existing trackers mainly face the problems of deep convolutional feature redundancy and a lack of positive samples in the target tracking process.In order to remove redundant features in deep convolutional networks,an attention mechanism model based on the fusion of the spatial domain and channel domain,which includes three continuous sub-modules:spatial self-attention,channel attention and spatial attention.On this basis,a deep convolutional network suitable for various vision algorithms is constructed.A feature enhancement technique based on feature modification is meant to extract target deep convolutional features in order to address the issues of negative feedback of redundant features and lack of positive samples.The experimental results show that in the mainstream residual networks(ResNet)object classification tasks,the addition of the feature modification module can significantly reduce the Top-1 and Top-5 error rates,and will not cause additional calculation or network structure adjustment,achieving lightweight insertion.Several tracking methods are combined with the feature modification module to improve tracking performance and address the discriminator's over-fitting issue.
As customer demand for multi-variety and small-batch production increases, dynamic disturbances place greater demands on manufacturing systems. To address such challenges, researchers proposed the multi-agent manufacturing system. However, conventional agent negotiation typically relies on pre-defined and fixed heuristic rules, which are ill-suited to managing complex and fluctuating disturbances. In current implementations, mainstream approaches based on reinforcement learning require the development of simulators and training models specific to a given shopfloor, necessitating substantial computational resources and lacking scalability. To overcome this limitation, the present study proposes a Large Language Model-based (LLM-based) multi-agent manufacturing system for intelligent shopfloor management. By defining the diverse modules of agents and their collaborative methods, this system facilitates the processing of all workpieces with minimal human intervention. The agents in this system consist of the Machine Server Module (MSM), Bid Inviter Module (BIM), Bidder Module (BM), Thinking Module (TM), and Decision Module (DM). By harnessing the reasoning capabilities of LLMs, these modules enable agents to dynamically analyze shopfloor information and select appropriate processing machines. The LLM-based modules, predefined by system prompts, provide dynamic functionality for the system without the need for pre-training. Extensive experiments were conducted in physical shopfloor settings. The results demonstrate that the proposed system exhibits strong adaptability, and achieves superior performance (makespan) and stability (as measured by sample standard deviation) compared to other approaches without requiring pre-training.
The multi-agent manufacturing system has emerged as a well-established paradigm in intelligent manufacturing. Presently, challenges such as limited adaptability, elevated maintenance expenses, and complexities in enabling local agent deployment at end devices persist. To address such issues, a deployment model for the multi-agent manufacturing system was proposed, leveraging a cloud-edge collaboration architecture. However, managing agents effectively in this environment to establish resilient services, which are services capable of maintaining high availability, stability, and reliability even in the face of uncertainty, emergencies, or failures, for manufacturing systems remains a critical challenge that requires immediate resolution. In the present study, a cloud-edge-end oriented deployment architecture for multi-agent manufacturing system was proposed, and a real-time mapping method between edge agents and production resources based on the 5th generation mobile communication technology is constructed. At the same time, a resource optimisation method called swarm avian evolutionary algorithm is proposed. This method integrates particle swarm optimisation and meta-heuristics to minimise computation time and enhance system response speed. Finally, the proposed resource optimisation method is compared with the genetic algorithm, particle swarm optimisation, and snake optimiser algorithms. The results demonstrate that the convergence time is significantly reduced, indicating that the proposed method offers superior performance.
To address the limitations of existing dual-attention vision-language models, including poor feature representation quality, inefficient attention interaction, and difficulties in lightweight deployment, this paper proposes an enhanced vision-language model that incorporates a frequency-domain dynamic filtering mechanism and a parallel attention interaction paradigm. The model features three core in-novations: First, a Wavelet Adaptive Dynamic Filtering module (WADT) is embedded before the dual-attention module to purify features, enhance fine details, and suppress background noise interference. Second, a dual-branch parallel architecture combined with a bidirectional cross-attention mechanism is adopted to reconstruct the attention interaction logic, solving the semantic disconnect problem of serial interaction and improving the fusion efficiency of cross-modal and cross-dimensional features. Third, lightweight components are constructed to balance detection performance and computational overhead, achieving a favorable trade-off between accuracy and complexity. Extensive pre-training is conducted on public datasets, followed by quantitative evaluation, qualitative validation, and ab-lation studies on mainstream vision tasks such as image classification, object detection, and visual grounding. Experimental results demonstrate that the proposed model achieves 59.2% mean Av-erage Precision (mAP) on object detection, outperforming state-of-the-art baselines by 4.2%. For zero-shot visual grounding, the Recall@1 metric reaches 85.2%, surpassing contemporary advanced models. Meanwhile, the model reduces parameters by 45%, realizing collaborative optimization of high performance and low overhead, and provides a feasible solution for the lightweight deployment of vision-language models.
In order to optimize the scale of the tracker's output,this paper focuses on how to fully utilize the semantic information of the target obtained by image semantic segmentation in complex scenes.It also designs an image semantic segmentation network based on the optimization of the attention mechanism to optimize the target tracker's output and the input of the features,which can realize plug-and-play for various algorithms.The image semantic segmentation mask is used to obtain the rotating frame boundary of the target,and the denoising optimization of the features in the input phase of the target is carried out according to the rotating and non-rotating frame boundaries of the target to attenuate the influence of the background noise on the discriminator of the tracker.The structure of the designed network,training,calibration of the target's rotating frame,and denoising of the tracker's input features are discussed,respectively.The correctness of the target motion model in resolving the scale calibration of the target during target tracking is verified through experimental comparison analysis on public datasets OTB100,VOT2016 and VOT2018.This enhances the accuracy and resilience of target tracking.
The prediction of pedestrian trajectories plays a crucial role in practical traffic scenarios. However, current methodologies have shortcomings, such as overlooking pedestrians' perception of motion information from neighbor groups, employing simplistic and fixed social state interaction models, and lacking in final position correction. To address these issues, SocialTrans is proposed. It utilizes global observations to model the motion states of pedestrians and their neighbors, constructing separate state tensors to encapsulate social interaction information between them. This design includes a Subject Intention Extraction Module and a Neighbor Perception Intentions Extraction Module, which operate in parallel throughout the observation period to facilitate deep interaction of social states rather than simple end-to-end external fusion. Furthermore, a trajectory prediction optimizer is developed to correct final position predictions and simulate pedestrian motion diversity through trajectory clustering. Experimental validation is conducted on the ETH/UCY and SDD public datasets to evaluate the effectiveness of the proposed approach. The results demonstrate the method's capability to learn historical trajectory information, achieve high-precision predictions, and achieve state-of-the-art performance, particularly outperforming existing SOTA models on the SDD dataset. The algorithm will be made available at https://github. com/XiaodZhao/SocialTrans.
Current autonomous driving systems often struggle to balance decision-making and motion control while ensuring safety and traffic rule compliance, especially in complex urban environments. Existing methods may fall short due to separate handling of these functionalities, leading to inefficiencies and safety compromises. To address these challenges, we introduce UDMC, an interpretable and unified Level 4 autonomous driving framework. UDMC integrates decision-making and motion control into a single optimal control problem (OCP), considering the dynamic interactions with surrounding vehicles, pedestrians, road lanes, and traffic signals. By employing innovative potential functions to model traffic participants and regulations, and incorporating a specialized motion prediction module, our framework enhances on-road safety and rule adherence. The integrated design allows for real-time execution of flexible maneuvers suited to diverse driving scenarios. High-fidelity simulations conducted in CARLA exemplify the framework's computational efficiency, robustness, and safety, resulting in superior driving performance when compared against various baseline models. Our open-source project is available at https://github.com/henryhcliu/udmc_carla.git.
This paper aims to improve the performance of video multimodal large language models (MLLM) via long and rich context (LRC) modeling. As a result, we develop a new version of InternVideo2.5 with a focus on enhancing the original MLLMs' ability to perceive fine-grained details and capture long-form temporal structure in videos. Specifically, our approach incorporates dense vision task annotations into MLLMs using direct preference optimization and develops compact spatiotemporal representations through adaptive hierarchical token compression. Experimental results demonstrate this unique design of LRC greatly improves the results of video MLLM in mainstream video understanding benchmarks (short long), enabling the MLLM to memorize significantly longer video inputs (at least 6x longer than the original), and master specialized vision capabilities like object tracking and segmentation. Our work highlights the importance of multimodal context richness (length and fineness) in empowering MLLM's innate abilites (focus and memory), providing new insights for future research on video MLLM. Code and models are available at https://github.com/OpenGVLab/InternVideo/tree/main/InternVideo2.5
Cloud Manufacturing (CMfg) is a new manufacturing mode that provides efficient manufacturing services to customers by centrally scheduling manufacturing resources distributed across various regions. In CMfg, each participant is an independent economic entity with distinct objectives and effectively achieving the objectives of customers, suppliers, and the CMfg platform under limited resources is a significant challenge. To solve this problem, this study first proposed a three-level multi-task optimization (TMTO) model. The upper-level and lower-level of the TMTO model respectively optimize the personalized objectives of customers and suppliers, as well as the objectives of the CMfg platform are optimized at the middle-level. Subsequently, a skill vector-guided multi-task optimization algorithm (SMTOA) is proposed to collaboratively optimize the objectives of all participants, with the skill vector designed to evaluate the ability of scheduling schemes to meet the objectives of all customers and suppliers. Finally, experimental cases based on an aerospace manufacturing enterprise confirm the effectiveness of the TMTO model and the advantages of SMTOA in solving the TMTO model.
Pedestrian 3D pose tracking in multi-view scenarios has extensive practical applications. However, existing methods often overlook the overall tracking accuracy of pedestrians, particularly the issues of missing and erroneous tracking caused by severe occlusions, disappearances, and reappearances. It further affects the accuracy of pose point association. To address these limitations, a two-stage method is proposed, involving tracking with an exceptionally low error rate, followed by obtaining higher precision 3D pose points. Firstly, a multi-object tracking model is introduced, which integrates feature association-validation-updating and employs dynamic thresholding strategy to achieve high-accuracy matching of multiple individuals in multi-view scenarios by computing similarity with feature pool templates. Additionally, a Gaussian Mixture-based feature pool updating model ensures the universality of stored features to solve the reappearance problem. Secondly, a pedestrian 2D pose detection and 3D pose reprojection method based on SMPL (Skinned Multi-Person Linear model) is proposed, which detects more complete pose points than OpenPose in complex scenes and better conforms to the distribution principles of human skeletal pose points. To validate the advancedness of the proposed method, the Shelf and Campus public datasets are re-annotated. Experimental results demonstrate the excellent performance of the proposed method in overall error control in complex environments, outperforming existing methods in multi-object tracking and pose point estimation accuracy and completeness.
At present, the production control process of a discrete manufacturing workshop is characterized by high concurrency, mixed production lines and difficulty in prediction, which lead to uncertainty caused by dynamic disturbances and challenges in production control. Traditional system architectures struggle to handle these uncertainties flexibly and adaptively. To address these issues, an adaptive production scheduling system for the workshop is proposed, utilizing the Multi-agent Cyber Physical System (CPS-MAS) framework. This system integrates self-organization mechanisms and self-adaptive decision-making mechanisms to achieve cooperative optimal control of manufacturing resources. Using multi-agent technology, the resource model in the information space is encapsulated into an intelligent Cyber Physical System (CPS)-Agent model with cognitive interaction and autonomous decision-making capabilities. The improved contract network protocol (CNP) is utilized to the constructed agent, enabling their collaboration and competition to support the self-organization, negotiation, and assignment of manufacturing tasks. Based on multi-agent real-time perception and interactive negotiation, an adaptive control model of the manufacturing process is constructed based on Proportion Integration Differentiation (PID) control principle. This model is trained with the multi-layer perceptron that integrates an attention mechanism. The production strategy and parameters of the agent cooperative network are dynamically adjusted to enable dynamic decision-making optimization under disturbances. The proposed method is verified by experiments in scenarios involving machine failure, emergency order insertion and due date changes, proving its effectiveness.
Image captioning is a cross-modal task that combines computer vision and natural language processing. The model is required to generate an appropriate caption for the given image. To address this challenge, we proposed a Residual Gated Transformer, RGFormer, as an enhancement of Transformer architecture. The model based on CLIP and RGFormer, CRM, is then proposed for image captioning. CRM utilizes the encoder of CLIP to extract image features as a prefix to the caption. The prefix is projected into language space using RGFormer, a lightweight mapping network, and then fed into GPT-2 to generate captions. CLIP was trained on an extensive dataset comprising image-text pairs, which contains rich visual and semantic information and is exceptionally well-suited for vision-language tasks. The core idea of CRM is to reduce the disparity between visual and textual representations by using RGFormer to accomplish the cross-modal task. CRM could generate meaningful captions for diverse and large-scale datasets in a short training time without additional annotations or pre-training. Quantitative evaluation experiments show that CRM achieves results comparable to some advanced models on the COCO Caption dataset more efficiently.
In complex tracking environments, existing trackers primarily encounter issues of redundant deep convolutional features and a shortage of positive samples in the target tracking process. To address these challenges, an attention mechanism model DFMM, the Deep Feature Modification Model, is proposed based on the fusion of spatial and channel domains. This model comprises three consecutive sub-modules: spatial self-attention, channel attention, and spatial attention. Building upon this, a deep convolutional network adaptable to various visual algorithms is constructed. Additionally, strategies for feature extraction and enhancement based on feature modification are designed to mitigate problems such as redundant feature negative feedback and a lack of positive samples. Experimental results demonstrate that integrating the feature modification module in mainstream ResNet target classification tasks significantly reduce Top-1 and Top-5 error rates without incurring additional computational overhead or necessitating network structure adjustments, achieving lightweight integration. Furthermore, incorporating the feature modification module in multiple related tracking algorithms enhances tracking performance and addresses discriminator overfitting issues.
Aiming at the problems that most of the existing methods of constructing 3D models of human body based on 2D human body surface pose points will lead to continuous modeling jitter and local distortion of the modeling results, we propose a 3D pose point detection method based on the skinned multiplayer linear model (SMPL) in the human body, which maps the 2D pose points of the body to 3D pose points of the real scene in the multiview perspective through a clustering algorithm, and introduces Kalman filtering to de-noise the human body pose points. The Kalman filter is introduced to denoise the human body posture points. In the process of constructing a 3D model of the human body based on 3D pose points, we construct an end-to-end human 3D modeling network (SMPL-VAE) based on the correction of gradient descent regression network by the automatic variational approach (VAE), which is more in line with the local modeling of the human body’s motion structure while maintaining the overall proportion.The results on open dataset Shelf show that our methods improve the quality of human post point detection and modeling.
In recent years, there has been a notable surge in investment interest in the establishment of Digital Twin (DT) shopfloors, underscoring the growing importance of DT technology. As such, there has been a heightened demand for the creation of DT models. Nevertheless, manual operations continue to play a significant role in the construction process, irrespective of time and financial considerations. Hence, it is imperative to develop an approach for the rapid construction of DT shopfloors aimed at reducing the reliance on manual operations. To this end, a point cloud based expeditious construction framework for DT shopfloors was proposed in the present study, which comprises pre-processing, processing, and post-processing modules. The pre-processing module is responsible for providing a clear point cloud to the processing module, and the post-processing module will reuse a historical DT model directly or establish a new one manually according to the result from the processing module. To achieve the task of determining whether a reusable model exists and which model to reuse for the post-processing module based on the input clear point cloud, the Transfer Point Cloud Net (TPCN) is employed as the processing network, comprising multiple enhanced Transformer blocks. Additionally, it demonstrates the capability to transfer knowledge derived from publicly available datasets, thereby diminishing the need for reliance on proprietary industrial data for training purposes. Moreover, this capacity can reduce the number of parameters necessary for training TPCN, leading to substantial time and computational resource savings. Comparative experiments were also conducted to validate the performance of the TPCN.
AI加速器在空间探索应用时需要考虑到空间辐射环境下SEE引发的软错误.在AI加速器设计过程中,需要对其SEE容错能力和可靠性进行评估,本文对Lenet-5的加速器进行了SEU故障注入,提出了一种从网络结构与电路模块映射的角度进行统计评估的方法.实验结果证明,在神经网络中,由于AI加速器计算数据大的特点,发生在权重和特征图的SEU错误在传播过程中有可能会被池化层屏蔽掉,SEU错误发生在靠近输出的层级比靠近输入的层级更容易导致识别准确率的下降.此外,实验还发现,在加速器电路模块映射中,负责产生使能信号和地址控制信号的控制单元CTRL比处理单元PE和存储单元MEM更容易被SEU错误所影响,严重时会影响加速器的正常运行.最后本文针对评估结果,进行了STMR加固措施对CTRL进行了加固,相比于FTMR,极大地减少了面积开销.
Future pedestrian trajectory prediction in first-person videos offers great prospects to help autonomous vehicles and social robots to enable better human-vehicle interactions. Given an egocentric video stream, we aim to predict the location and depth (distance between the observed person and the camera) of his/her neighbors in future frames. To locate their future trajectories, we mainly consider three main factors: a) It is necessary to restore the spatial distribution of pedestrians in 2D image to 3D space, i.e., to extract the distance between the pedestrian and the camera which is often neglected. b) It is critical to utilize neighbors' poses to recognize their intentions. c) It is important to learn human-vehicle interactions from the pedestrian's historical trajecto-ries. We propose to incorporate these three factors into a multi-channel tensor to represent the main features in real-life 3D space. We then put this tensor into an innovative end-to-end fully convolutional network based on transformer architecture. Experimental results reveal our method outperforms other state-of-the-art methods on public benchmarks MOT15, MOT16 and MOT17. The proposed method will be useful to understand human -vehicle interaction and helpful for pedestrian collision avoidance.(c) 2023 Elsevier B.V. All rights reserved.