Reliable communication and sensing are essential for automated control in customized Flexible Manufacturing Systems (FMS), where AGV networks frequently encounter Non Line-of-Sight (NLoS) conditions and interference from dense layouts and metallic obstacles. These propagation effects degrade both communication and sensing reliability, while conventional industrial networks further suffer from separate signal processing that delays perception–control feedback and leads to inefficient network utilization in dynamic environments. Integrated Sensing and Communication (ISAC) offers joint perception and communication to enable real-time, reliability-critical AGV coordination. However, current reliability analyses of ISAC systems often assume static channels and topologies, overlooking the coupled dynamics of AGV mobility, control logic, and volatile Line-of-Sight (LoS) and NLoS conditions. To address these issues, we develop an agent-based ISAC reliability analysis approach that couples mobility and channel dynamics for reliability assessment and infrastructure optimization under control-driven and performance-enhanced schemes. Simulation results reveal trade offs between deterministic control and reliability characteristics, where the performance-enhanced zero-forcing (PE-ZF) strategy achieves higher SINR and success probability. A weighted reliability-based infrastructure optimization is further proposed to enhance coverage and robustness, validating the effectiveness of the agent-based ISAC framework in flexible manufacturing environments.
Large language models (LLMs) have recently demonstrated remarkable success in mathematical reasoning. Despite progress in methods like chain-of-thought prompting and self-consistency sampling, these advances often focus on final correctness without ensuring that the underlying reasoning process is coherent and reliable. This paper introduces Step-KTO, a training framework that combines process-level and outcome-level binary feedback to guide LLMs toward more trustworthy reasoning trajectories. By providing binary evaluations for both the intermediate reasoning steps and the final answer, Step-KTO encourages the model to adhere to logical progressions rather than relying on superficial shortcuts. Our experiments on challenging mathematical benchmarks show that Step-KTO significantly improves both final answer accuracy and the quality of intermediate reasoning steps. For example, on the MATH-500 dataset, Step-KTO achieves a notable improvement in Pass@1 accuracy over strong baselines. These results highlight the promise of integrating stepwise process feedback into LLM training, paving the way toward more interpretable and dependable reasoning capabilities.
Hallucination, the generation of factually incorrect information, remains a significant challenge for large language models (LLMs), especially in open-domain long-form generation. Existing approaches for detecting hallucination in long-form tasks either focus on limited domains or rely heavily on external fact-checking tools, which may not always be available. In this work, we systematically investigate reference-free hallucination detection in open-domain long-form responses. Our findings reveal that internal states (e.g., model's output probability and entropy) alone are insufficient for reliably (i.e., better than random guessing) distinguishing between factual and hallucinated content. To enhance detection, we explore various existing approaches, including prompting-based methods, probing, and fine-tuning, with fine-tuning proving the most effective. To further improve the accuracy, we introduce a new paradigm, named RATE-FT, that augments fine-tuning with an auxiliary task for the model to jointly learn with the main task of hallucination detection. With extensive experiments and analysis using a variety of model families & datasets, we demonstrate the effectiveness and generalizability of our method, e.g., +3% over general fine-tuning methods on LongFact.
Solving mathematics problems has been an intriguing capability of language models, and many efforts have been made to improve reasoning by extending reasoning length, such as through selfcorrection and extensive long chain-of-thoughts. While promising in problem-solving, advanced long reasoning chain models exhibit an undesired uni-modal behavior, where trivial questions require unnecessarily tedious long chains of thought. In this work, we propose a way to allow models to be aware of inference budgets by formulating it as utility maximization with respect to an inference budget constraint, hence naming our algorithm Inference Budget-Constrained Policy Optimization (IBPO). In a nutshell, models fine-tuned through IBPO learn to "understand" the difficulty of queries and allocate inference budgets to harder ones. With different inference budgets, our best models are able to have a 4.14% and 5.74% absolute improvement (8.08% and 11.2% relative) on MATH500 using 2.16x and 4.32x inference budgets respectively, relative to LLaMA3.1 8B Instruct. These improvements are approximately 2x those of self-consistency under the same budgets.
In recent years, sequence prediction, particularly in natural language processing tasks, has made significant progress due to advanced neural network architectures like Transformer and enhanced computing power. However, challenges persist in modeling and analyzing certain types of sequence data, such as human daily activities and competitive ball games. These segmented sequence data are characterized by short length, varying local dependencies, and coarse-grained unit states. These characteristics limit the effectiveness of conventional probabilistic graphical models and attention-based or recurrent neural networks in modeling and analyzing segmented sequence data. To address this gap, we introduce a novel generative model for segmented sequences, employing an ensemble of multiple variable-order Markov models (VOMMs) to flexibly represent state transition dependencies. Our approach integrates probabilistic graphical models with neural networks, surpassing the representation capabilities of single high-order or variable-order Markov models. Compared to end-to-end deep learning models, our method offers improved interpretability and reduces overfitting in short segments. We demonstrate the efficacy of our proposed method in two tasks: predicting tennis shot types and forecasting daily action sequences. These applications highlight the broad applicability of our segmented sequence modeling approach across diverse domains.
Recent advances in large language models (LLMs) have demonstrated significant progress in performing complex tasks. While Reinforcement Learning from Human Feedback (RLHF) has been effective in aligning LLMs with human preferences, it is susceptible to spurious correlations in reward modeling. Consequently, it often introduces biases-such as length bias, sycophancy, conceptual bias, and discrimination that hinder the model's ability to capture true causal relationships. To address this, we propose a novel causal reward modeling approach that integrates causal inference to mitigate these spurious correlations. Our method enforces counterfactual invariance, ensuring reward predictions remain consistent when irrelevant variables are altered. Through experiments on both synthetic and real-world datasets, we show that our approach mitigates various types of spurious correlations effectively, resulting in more reliable and fair alignment of LLMs with human preferences. As a drop-in enhancement to the existing RLHF workflow, our causal reward modeling provides a practical way to improve the trustworthiness and fairness of LLM finetuning.
Large Language Models (LLMs) have demonstrated impressive capabilities in various tasks, including instruction following, which is crucial for aligning model outputs with user expectations. However, evaluating LLMs' ability to follow instructions remains challenging due to the complexity and subjectivity of human language. Current benchmarks primarily focus on single-turn, monolingual instructions, which do not adequately reflect the complexities of real-world applications that require handling multi-turn and multilingual interactions. To address this gap, we introduce Multi-IF, a new benchmark designed to assess LLMs' proficiency in following multi-turn and multilingual instructions. Multi-IF, which utilizes a hybrid framework combining LLM and human annotators, expands upon the IFEval by incorporating multi-turn sequences and translating the English prompts into another 7 languages, resulting in a dataset of 4,501 multilingual conversations, where each has three turns. Our evaluation of 14 state-of-the-art LLMs on Multi-IF reveals that it presents a significantly more challenging task than existing benchmarks. All the models tested showed a higher rate of failure in executing instructions correctly with each additional turn. For example, o1-preview drops from 0.877 at the first turn to 0.707 at the third turn in terms of average accuracy over all languages. Moreover, languages with non-Latin scripts (Hindi, Russian, and Chinese) generally exhibit higher error rates, suggesting potential limitations in the models' multilingual capabilities. We release Multi-IF prompts and the evaluation code base to encourage further research in this critical area.
This paper studies a novel quality-oriented efficient distributed framework for nonlinear plant-wide industrial quality-related process monitoring. In this strategy, process variables contained in the local unit are divided into quality-related and quality-unrelated parts using the elastic network. Then, the least absolute shrinkage and selection operator technique is utilized to select the communication variables that are highly relevant to the quality-related part of the local unit from neighboring units, which not only improves the quality-oriented process monitoring performance but also reduces redundant communications. Then, for the reorganized quality-related part of the local unit, a reasonable orthogonal decomposition is developed to cope with the inherent flaws of kernel partial least squares. This decomposition further divides the process variable space into two orthogonal parts. For the remaining quality-unrelated part of the local unit, the kernel principal component analysis with a combined index is used to monitor it. Finally, the Bayesian fusion is used to improve the monitoring efficiency. The proposed scheme and the existing methods are compared using the Tennessee Eastman benchmark process, demonstrating the superiority and effectiveness of the proposed method.Note to Practitioners-For plant-wide process monitoring, a novel quality-oriented efficient distributed monitoring strategy is developed in this paper, which not only considers the monitoring of the quality variables within systems but also emphasizes the communication efficiency between local units and neighboring units. By using the proposed strategy, local unit and neighboring unit variables can be initially filtered by applying a combination of the elastic network and the least absolute shrinkage and selection operator technique, which not only takes into account the correlation between local units and neighboring units but also reduces unnecessary information transfer from neighboring units. As a result, it improves the efficiency of distributed monitoring and ensures the accuracy of quality-related fault detection. Furthermore, the supervised process monitoring for quality variables is realized with the help of the proposed strategy. By utilizing the monitoring results, practitioners can accurately determine whether the fault type is quality-related or quality-unrelated. This information facilitates the design of a more targeted fault-tolerant control scheme, reducing unnecessary fault-tolerant control actions and enhancing the efficient utilization of the control system, ultimately leading to energy savings. Finally, the incorporation of the Bayesian fusion strategy enables the generation of both global fault and local fault detection indicators. This feature proves beneficial for designing subsequent visualization platforms, providing comprehensive information for fault analysis and system visualization.
Unmanned Aerial Vehicle (UAV) ad hoc network has achieved significant growth for its flexibility, extensibility, and high deployability in recent years. The application of clustering scheme for UAV ad hoc network is imperative to enhance the performance of throughput and energy efficiency. In conventional clustering scheme, a single cluster head (CH) is always assigned in each cluster. However, this method has some weaknesses such as overload and premature death of CH when the number of UAVs increased. In order to solve this problem, we propose a dual-cluster-head based medium access control (DCHMAC) scheme for large-scale UAV networks. In DCHMAC, two CHs are elected to manage resource allocation and data forwarding cooperatively. Specifically, two CHs work on different channels. One of CH is used for intra-cluster communication and the other one is for inter-cluster communication. A Markov chain model is developed to analyse the throughput of the network. Simulation result shows that compared with FM-MAC (flying ad hoc networks multi-channel MAC, FM-MAC), DCHMAC improves the throughput by approximately 20%∼50% and prolongs the network lifetime by approximately 40%.
Timely and accurate monitoring of power line safety is a key step in ensuring urban production and daily life. The rapid development of vehicle-borne mobile laser scanning (MLS) technology has provided an effective solution for intelligent maintenance of power grids. This paper proposes a comprehensive method for accurate extraction, missing completion, and intrusion risk detection of urban power lines. Firstly, local linear geometric features are described using principal component analysis (PCA) covariance matrix, and outliers are accurately separated using a joint Bezier curve distance threshold. Secondly, the statistical inference function estimation method is employed to accurately simulate the complete topology of power lines. Finally, a fast intrusion risk detection method based on the separating axis theorem (SAT) for boundary box neighboring objects is developed. Experimental results demonstrate that our method achieves a recall rate and precision rate of 98.28 % and 98.71 %, respectively, for power line extraction. The root mean square error (RMSE) accuracy and coordinate residual of catenary line fitting are both below 0.04 m, indicating strong noise resistance and robustness. Compared to mainstream geometric extraction methods, our approach exhibits superior performance.
Reinforcement learning from human feedback (RLHF) has become the leading approach for fine-tuning large language models (LLM). However, RLHF has limitations in multi-task learning (MTL) due to challenges of reward hacking and extreme multi-objective optimization (i.e., trade-off of multiple and/or sometimes conflicting objectives). Applying RLHF for MTL currently requires careful tuning of the weights for reward model and data combinations. This is often done via human intuition and does not generalize. In this work, we introduce a novel post-training paradigm which we called Constrained Generative Policy Optimization (CGPO). The core of CGPO is Mixture of Judges (MoJ) with cost-efficient constrained policy optimization with stratification, which can identify the perfect blend in RLHF in a principled manner. It shows strong empirical results with theoretical guarantees, does not require extensive hyper-parameter tuning, and is plug-and-play in common post-training pipelines. Together, this can detect and mitigate reward hacking behaviors while reaching a pareto-optimal point across an extremely large number of objectives. Our empirical evaluations demonstrate that CGPO significantly outperforms standard RLHF algorithms like PPO and DPO across various tasks including general chat, STEM questions, instruction following, and coding. Specifically, CGPO shows improvements of 7.4 Arena-Hard (STEM reasoning), and consistent gains in other domains like math and coding. Notably, PPO, while commonly used, is prone to severe reward hacking in popular coding benchmarks, which CGPO successfully addresses. This breakthrough in RLHF not only tackles reward hacking and extreme multi-objective optimization challenges but also advances the state-of-the-art in aligning general-purpose LLMs for diverse applications.
In consideration of the significant changes in urban environmental scenes and the impact of factors such as signal obstruction on positioning, the spatial positions of multi-temporal Mobile Laser Scanning (MLS) point clouds collected in the same area are inconsistent, which makes it challenging to provide a complete expression of the urban road environment. To address this issue, a precise registration algorithm based on pole-like targets is proposed. The core steps include registration primitive extraction: a cylindrical spatial neighborhood is established based on the distinctive key points, and precise extraction of pole-like geospatial point cloud data is performed using the Euclidean clustering method with the normal curvature constraint. Achieving optimal spatial transformation: a circle center fitting method with additional Mean Absolute Error (MAE) loss function constraints was developed based on the stable elemental unit which was utilized to accurately construct homonymous feature point pairs. Validation results from multiple real urban point cloud datasets collected from different experimental areas and platforms demonstrate that our method achieves superior registration performance, with mean residuals below 5 cm in both the planar and elevation directions. The proposed method exhibits strong universality and registration robustness for multi-temporal point cloud data of the homologous or cross-sources.
We present a series of long-context LLMs that support effective context windows of up to 32,768 tokens. Our model series are built through continual pretraining from Llama 2 with longer training sequences and on a dataset where long texts are upsampled. We perform extensive evaluation on language modeling, synthetic context probing tasks, and a wide range of research benchmarks. On research benchmarks, our models achieve consistent improvements on most regular tasks and significant improvements on long-context tasks over Llama 2. Notably, with a cost-effective instruction tuning procedure that does not require human-annotated long instruction data, the 70B variant can already surpass gpt-3.5-turbo-16k's overall performance on a suite of long-context tasks. Alongside these results, we provide an in-depth analysis on the individual components of our method. We delve into Llama's position encodings and discuss its limitation in modeling long dependencies. We also examine the impact of various design choices in the pretraining process, including the data mix and the training curriculum of sequence lengths -- our ablation experiments suggest that having abundant long texts in the pretrain dataset is not the key to achieving strong performance, and we empirically verify that long context continual pretraining is more efficient and similarly effective compared to pretraining from scratch with long sequences.
The increasing complexity of urban environments introduces additional uncertainty to the deployment of the autonomous vehicular network. A novel road infrastructure cooperative detection model using Joint Communication and Sensing (JCS) technology is proposed in this article to simultaneously achieve high-efficient communication and obstacle detection for urban autonomous vehicles. To suppress the performance fluctuation caused by shadowing and obstruction to the JCS signals, we first derive the statistic of road obstacles from the Geographic Information System (GIS). Then, the analysis of JCS channel characteristics and shadowing factors are presented using Line-of-Sight and Non-Line-of-Sight (LoS and NLoS) channel models under the complex urban scenario. A stochastic geometry approach is applied to analyze the interference factors and the probability distribution of successful JCS detection and communication. Simulations have been made to verify the cooperative detection model by probability analysis based on LoS and NLoS channels, and the numerical results demonstrate several different optimization methods for the deployment of JCS road infrastructures. Finally, we simulated and analyzed a deployment optimization method for JCS road infrastructures that complied with the standard of urban traffic-spot structure placement.
Quality-related process monitoring as a supervised technology has increasingly attracted attention in complex industries. Various approaches have been studied to cope with this issue. Nevertheless, these methods cannot reasonably decompose the process variable space, resulting in deficiencies in monitoring quality-related faults. To handle this issue, this paper presents an orthogonal kernel partial least squares improved kernel least squares with a preprocessing-modeling-postprocessing (PMP) structure to implement quality-related process monitoring with more proper decomposition and more straightforward monitoring logic. Compared with the previous approaches, a nonlinear preprocessing technology is presented to eliminate the quality-unrelated knowledge of process variables, enormously enhancing the interpretability of modeling and improving the monitoring efficiency. Then, a proper decomposition is presented to decompose the kernel matrix into two orthogonal parts, significantly improving the monitoring performance. The theoretical analysis of the proposed method is provided in this paper. Finally, two cases indicate the validity and superiority of the proposed method.
Kernel partial least squares (KPLS) has poor robustness and cannot achieve effective monitoring for key performance indicators (KPI). This study investigates a new KPI-oriented robust KPLS approach to mitigate these drawbacks. In this methodology, the robust KPLS inspired by the gradient boosting principle incorporates a weighting matrix into the KPLS to mitigate the influence of outliers. Simultaneously, the associated coefficient matrix is derived in detail. Then, a decomposition approach is used to separate the process variable space into two orthogonal parts. Two strategies are discussed to obtain the unknown projection matrices based on the kernel principal component analysis and the elastic network frameworks. Finally, the performance of the proposed methods in terms of prediction, monitoring, and robustness to outliers is evaluated by the Tennessee Eastman process and the three-phase flow facility. The results show that the proposed methods have good prediction accuracy and monitoring performance in the presence of outliers, demonstrating the effectiveness and advantages of the proposed approaches.
Vehicle-borne Mobile Laser Scanning (MLS) point cloud is one of the key components of 3D spatial information, which provides important geographic data support for the development of smart cities and autonomous driving. However, urban high-rise buildings will obscure the Global Navigation Satellite System (GNSS) positioning signal of the vehicle-borne MLS system, causing the problem of the inconsistent location of revisited point clouds collected in the same area, and the existing methods are still limited by low accuracy and efficiency when registering large-scale point clouds. Aiming at the complexity and timeliness of vehicle-borne point cloud registration, this paper proposes an improved Iterative Closest Point (ICP) point cloud accurate registration method based on road markings. Firstly, the road feature image is generated based on the laser reflection intensity of the ground point cloud, and the Pix2Pix_L1 (P2P_L1) transformation model is introduced to realize the identification and classification of markings. Secondly, we take road classification markings as the registration primitives, a Random Sample Consensus (RANSAC) point cloud coarse registration method based on a Fast Point Feature Histogram (FPFH) descriptor is designed, and the Chi-square distance function is used to improve the stability of corresponding points. Then, according to the normal vector characteristics of marking point clouds, the angle between the normal vectors is constrained to extract edge points of markings and select effective corresponding closest point pairs, and the translation vector is preferentially calculated according to the offset characteristics of point cloud to optimize spatial transformation matrix. Finally, a linear interpolation method is developed to correct the pose of global point clouds. Experimental results show that our method achieves high registration accuracy and computational speed on multiple datasets, and has stronger performance than the most widely used geometric feature registration methods.
Automatic question generation (AQG) is the task of generating a question from a given passage and an answer. Most existing AQG methods aim at encoding the passage and the answer to generate the question. However, limited work has focused on modeling the correlation between the target answer and the generated question. Moreover, unseen or rare word generation has not been studied in previous works. In this paper, we propose a novel approach which incorporates question generation with its dual problem, question answering, into a unified primal-dual framework. Specifically, the question generation component consists of an encoder that jointly encodes the answer with the passage, and a decoder that produces the question. The question answering component then re-asks the generated question on the passage to ensure that the target answer is obtained. We further introduce a knowledge distillation module to improve the model generalization ability. We conduct an extensive set of experiments on SQuAD and HotpotQA benchmarks. Experimental results demonstrate the superior performance of the proposed approach over several state-of-the-art methods.
Prompt tuning is a new, efficient NLP transfer learning paradigm that adds a task-specific prompt in each input instance during the model training stage. It freezes the pre-trained language model and only optimizes a few task-specific prompts. In this paper, we propose a conditional prompt generation method to generate prompts for each input instance, referred to as the Instance-Dependent Prompt Generation (IDPG). Unlike traditional prompt tuning methods that use a fixed prompt, IDPG introduces a lightweight and trainable component to generate prompts based on each input sentence. Extensive experiments on ten natural language understanding (NLU) tasks show that the proposed strategy consistently outperforms various prompt tuning baselines and is on par with other efficient transfer learning methods such as Compacter while tuning far fewer model parameters.
Jie Tang (唐杰)合作论文数Department of Computer Science and Technology, Tsinghua University8