
惠普(Hewlett-Packard,简称HP)是信息科技(IT)公司之一,成立于1939年,总部位于美国加利福尼亚州帕洛阿尔托市。惠普下设三大业务集团:信息产品集团、打印及成像系统集团和企业计算机专业服务集团。 中国惠普有限公司总部位于北京,在上海、广州、沈阳、南京、西安、武汉、成都、深圳等都设有分公司。中国惠普在大连设有惠普全球呼叫中心,在重庆设有生产工厂,在天津设有数据中心。
Large Language Model (LLM) unlearning aims to remove targeted knowledge from a trained model, but practical deployments often require post-training quantization (PTQ) for efficient inference. However, aggressive low-bit PTQ can mask or erase unlearning updates, causing quantized models to revert to pre-unlearning behavior. We show that standard full-parameter fine-tuning often induce parameter changes that are too small to survive 4-bit quantization. We propose quantization-robust unlearning via low-rank adaptation (LoRA): we freeze the base model and concentrate unlearning into trainable adapters so that the effective update is preserved after quantization. On Llama-2-7B evaluated with MUSE dataset (BOOKS and NEWS), LoRA improves 4-bit utility by up to 7.93 points (NPO+GDR on BOOKS: 50.17 to 58.10) and yields higher 4-bit utility on NEWS for GA+GDR (40.06 to 44.82, increase of 4.76). LoRA also substantially reduces privacy leakage under 4-bit PTQ, e.g., for GA+KLR on BOOKS, PrivLeak moves from -25.68 to -5.86 (closer to ideal 0), while maintaining strong forgetting (VerMem and KnowMem near 0). Thus, using LoRA for Machine Unlearning is beneficial for scenarios where quantization is necessary for model deployment.
Direct alignment algorithms such as Direct Preference Optimization (DPO) fine-tune models based on preference data, using only supervised learning instead of two-stage reinforcement learning with human feedback (RLHF). We show that DPO encodes a statistical estimation problem over reward functions induced by a parametric policy class. When the true reward function that generates preferences cannot be realized via the policy class, DPO becomes misspecified, resulting in failure modes such as preference order reversal, worsening of policy reward, and high sensitivity to the input preference data distribution. On the other hand, we study the local behavior of two-stage RLHF for a parametric class and relate it to a natural gradient step in policy space. Our fine-grained geometric characterization allows us to propose AuxDPO, which introduces additional auxiliary variables in the DPO loss function to help move towards the RLHF solution in a principled manner and mitigate the misspecification in DPO. We empirically demonstrate the superior performance of AuxDPO on didactic bandit settings as well as LLM alignment tasks.
Cancer treatment is at the core a sequential decision-making problem with partial observability, latent patient heterogeneity, and explicit constraints on the budget for medical measurements. Unlike standard Reinforcement Learning (RL) approaches that control state trajectories, cancer treatments permanently modify patients' transition dynamics, changing how states evolve over time. We model cancer treatment as a belief-space planning problem using active inference, deriving an expected free-energy objective that unifies goal-directed control and information acquisition under measurement budgets without. We implement this framework using real clinical cancer data from the AACR Project GENIE Biopharma Collaborative dataset. Results on clinical data demonstrate a simultaneous patient categorization and high treatment efficacy, under real measurement and treatment constraints.
Recent work on Neural Network-based methods for nonlinear control use Lyapunov Functions to obtain controllers with guarantees of stability. However, Lyapunov-based methods are fundamentally limited: they cannot be used for smooth blending with formal Region of Attraction (RoA) expansion guarantees, and also fail to certify stability when unstable equilibria or saddle points are present. Density functions provide an alternate stability certificate, and address these limitations by certifying almost everywhere stability, and enable smooth blending of controllers. Learning valid density certificates is challenging due to integrability constraints, and the effect of density-based blending controllers on RoAs is not well understood. In this work, we provide the first guarantee that controllers blended with density functions yield RoAs containing the union of the RoAs achieved by the constituent controllers. Then, we propose a novel exponential characterization of density functions that provably satisfies the integrability condition, and introduce Neural Control Density Functions (NCDFs), that leverage this new parameterization. We also extend NCDFs for synthesizing safe-stable controllers by combining NCDFs with control barrier functions (NCDF-CBFs). Our experiments show that blended controllers obtain superior RoAs to state-of-the-art methods like Neural Lyapunov Control and Sum-of-Squares based techniques.
Abstract Casing wear prediction is critical in High Pressure, High Temperature (HPHT) wells, where low wear tolerance, driven by burst and collapse load scenarios and requiring thick-walled Casing (CSG), constrains operational flexibility and lifecycle well operation. This study estimates wear on a 14″ × 13-5/8″ CSG string based on actual drilling operations by reconstructing rotating exposure in a high-angle trajectory with a shallow Kick-Off Point (KOP), computing mechanical work, selecting section-specific wear factors, and deriving the post-drilling wear profile. Final wear values are incorporated into the as-built design to define the well operating envelope and ensure well integrity during operation. High-resolution time-based drilling data was used to rebuild the operational sequence of each Bottom-Hole Assembly (BHA) run. Depth-distributed revolutions were recalculated, enabling accurate mapping of real rotating exposure. Torque & Drag & Buckling (T&D&B) modelling was applied to estimate side forces and compute the mechanical work imparted to the CSG string. Wear factors were determined through an established model calibrated on North Sea datasets, differentiating between early-groove and progressed-groove conditions. The resulting work-based wear distribution was compared with pre-drill wear estimates to evaluate the impact of operational deviations. Actual rotating exposure significantly exceeded pre-drill assumptions, especially in the later hole section, where unplanned rotating operations more than doubled the expected revolutions. This increased rotational contact resulted in approximately a 9% rise in maximum predicted wear relative to conservative pre-drill cases, with peak wear reaching 13.7% of nominal wall thickness. The wear apex also shifted deeper than anticipated, highlighting the influence of operational realisations on both wear magnitude and groove location. The post-drilling wear results were obtained using the work-based methodology, rather than an update of the pre-drill model. These observations reinforce that considering revolution counts alone does not capture the true mechanical interaction between the drill string and CSG, as they overlook variations in side forces, Rate of Penetration (ROP), and localised contact conditions. In contrast, the work-based reconstruction method, grounded in depth-distributed mechanical work, delivers a more accurate representation of casing wear severity and distribution. The findings emphasise the need for high-resolution operational data, improved wear factor calibration, and systematic post-drilling validation to enhance CSG integrity management throughout well construction. This study introduces a field-ready workflow that replaces traditional revolution-based casing wear estimates with a work-based approach derived from reconstructed operations. By accounting for unplanned rotation, variable ROP, and section-specific wear progression, it yields more realistic wear predictions and provides a replicable methodology that can significantly strengthen casing wear management in complex HPHT wells.